Episode

How SRE Teams Use Infrastructure as Code to Prevent Configuration Drift

Podcast
The Site Reliability Podcast with Fexingo: SRE, Uptime, and Production Engineering
Published
Jun 23, 2026
Duration seconds
663
Processing state
not_requested
Canonical source
https://audio.fexingo.com/business/the-site-reliability-podcast/episode-0068.mp3
Audio
https://audio.fexingo.com/business/the-site-reliability-podcast/episode-0068.mp3
JSON
/v1/public/podcasts/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/episodes/how-sre-teams-use-infrastructure-as-code-to-prevent-configuration-drift
Markdown
/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-infrastructure-as-code-to-prevent-configuration-drift.md

Actions

  • POST https://stenobird.com/v1/public/podcasts/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/episodes/how-sre-teams-use-infrastructure-as-code-to-prevent-configuration-drift/transcription-requests
    Idempotently request low-priority transcript generation for this episode.
  • GET https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-infrastructure-as-code-to-prevent-configuration-drift.md
    Read the agent-friendly Markdown representation of this episode resource.

Summary

In Episode 68 of The Site Reliability Podcast, Lucas and Luna explore how SRE teams use infrastructure as code (IaC) to prevent configuration drift — the silent killer of production reliability. They break down a real incident at a mid-sized fintech company where a manual SSH change caused a partial outage, and how the team rebuilt their entire environment with Terraform and automated compliance checks. Lucas explains the technical difference between declarative and imperative IaC, and why treating infrastructure like application code — with pull requests, code reviews, and unit tests — reduces human error by over 80 percent. Luna shares a surprising stat from a 2025 Cloud Posse survey: 63 percent of SRE teams still rely on partial manual config changes. They discuss practical steps to start small, like using Terraform for a single service, and the importance of integrating IaC with incident response playbooks. If you're an SRE or platform engineer struggling with snowflake servers, this episode offers concrete strategies to enforce consistency and reduce toil. No fluff, just real ops lessons. #InfrastructureAsCode #ConfigurationDrift #Terraform #SiteReliabilityEngineering #DevOps #Automation #CloudComputing #IncidentResponse #Fintech #PlatformEngineering #SRE #ToilReduction #ComplianceAsCode #DeclarativeIaC #GitOps #Technology #FexingoBusiness #BusinessPodcast Keep every episode free: buymeacoffee.com/fexingo