{"podcast":{"title":"The Site Reliability Podcast with Fexingo: SRE, Uptime, and Production Engineering","slug":"the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923","podcast_index_feed_id":7871923,"rss_url":"https://feeds.fexingo.com/business/the-site-reliability-podcast.xml","website_url":"https://www.fexingo.com/","image_url":"https://audio.fexingo.com/business/the-site-reliability-podcast/cover.png","author":"Fexingo","episode_count":119,"summary":"Lucas and Luna cut through the noise around site reliability engineering to examine how real-world SRE teams balance uptime, incident response, and production change. Each episode takes a single concept — error budgets, toil automation, postmortem culture, capacity planning — and grounds it in a specific case: how a major streaming service reduced paging noise, how a payments platform rebuilt its incident command structure, or how a cloud provider manages multi-region failover. Lucas brings the numbers — latency percentiles, MTTR trends, SLO burn rates — while Luna pushes on the human and organizational trade-offs: What does a junior SRE need to know about on-call? How do you measure reliability without crushing innovation? Why do some blameless postmortems actually work? Together they treat SRE not as a certification topic but as a living practice, citing real outages, open-source tools, and engineering blogs. This show is for engineers, ops leads, and platform teams who already know the basics and want to debate the hard edges: Is 99.999% uptime always worth the cost? When should you deliberately degrade service to improve reliability? How do you design for resilience when your…","last_synced_at":"2026-07-18T22:18:41.635042+00:00","page_url":"https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923"},"episode":{"title":"How SRE Teams Use Infrastructure as Code to Prevent Configuration Drift","slug":"how-sre-teams-use-infrastructure-as-code-to-prevent-configuration-drift","published_at":"2026-06-23T10:16:36+00:00","page_url":"https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-infrastructure-as-code-to-prevent-configuration-drift","show_page_url":"https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923","url":"https://audio.fexingo.com/business/the-site-reliability-podcast/episode-0068.mp3","audio_url":"https://audio.fexingo.com/business/the-site-reliability-podcast/episode-0068.mp3","summary":"In Episode 68 of The Site Reliability Podcast, Lucas and Luna explore how SRE teams use infrastructure as code (IaC) to prevent configuration drift — the silent killer of production reliability. They break down a real incident at a mid-sized fintech company where a manual SSH change caused a partial outage, and how the team rebuilt their entire environment with Terraform and automated compliance checks. Lucas explains the technical difference between declarative and imperative IaC, and why treating infrastructure like application code — with pull requests, code reviews, and unit tests — reduces human error by over 80 percent. Luna shares a surprising stat from a 2025 Cloud Posse survey: 63 percent of SRE teams still rely on partial manual config changes. They discuss practical steps to start small, like using Terraform for a single service, and the importance of integrating IaC with incident response playbooks. If you're an SRE or platform engineer struggling with snowflake servers, this episode offers concrete strategies to enforce consistency and reduce toil. No fluff, just real ops lessons. #InfrastructureAsCode #ConfigurationDrift #Terraform #SiteReliabilityEngineering #DevOps #Automation #CloudComputing #IncidentResponse #Fintech #PlatformEngineering #SRE #ToilReduction #ComplianceAsCode #DeclarativeIaC #GitOps #Technology #FexingoBusiness #BusinessPodcast Keep every episode free: buymeacoffee.com/fexingo","meta_description":"In Episode 68 of The Site Reliability Podcast, Lucas and Luna explore how SRE teams use infrastructure as code (IaC) to prevent configuration drift — the…","key_points":[],"chapters":[],"topics":[],"duration_seconds":663,"processing_state":"not_requested","actions":[{"name":"request_transcript","method":"POST","url":"https://stenobird.com/v1/public/podcasts/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/episodes/how-sre-teams-use-infrastructure-as-code-to-prevent-configuration-drift/transcription-requests","description":"Idempotently request low-priority transcript generation for this episode."},{"name":"read_markdown","method":"GET","url":"https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-infrastructure-as-code-to-prevent-configuration-drift.md","description":"Read the agent-friendly Markdown representation of this episode resource."}]}}