{"podcast":{"title":"The Site Reliability Podcast with Fexingo: SRE, Uptime, and Production Engineering","slug":"the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923","podcast_index_feed_id":7871923,"rss_url":"https://feeds.fexingo.com/business/the-site-reliability-podcast.xml","website_url":"https://www.fexingo.com/","image_url":"https://audio.fexingo.com/business/the-site-reliability-podcast/cover.png","author":"Fexingo","episode_count":119,"summary":"Lucas and Luna cut through the noise around site reliability engineering to examine how real-world SRE teams balance uptime, incident response, and production change. Each episode takes a single concept — error budgets, toil automation, postmortem culture, capacity planning — and grounds it in a specific case: how a major streaming service reduced paging noise, how a payments platform rebuilt its incident command structure, or how a cloud provider manages multi-region failover. Lucas brings the numbers — latency percentiles, MTTR trends, SLO burn rates — while Luna pushes on the human and organizational trade-offs: What does a junior SRE need to know about on-call? How do you measure reliability without crushing innovation? Why do some blameless postmortems actually work? Together they treat SRE not as a certification topic but as a living practice, citing real outages, open-source tools, and engineering blogs. This show is for engineers, ops leads, and platform teams who already know the basics and want to debate the hard edges: Is 99.999% uptime always worth the cost? When should you deliberately degrade service to improve reliability? How do you design for resilience when your…","last_synced_at":"2026-07-18T22:18:41.635042+00:00","page_url":"https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923"},"episode":{"title":"How SRE Teams Use Runbooks to Standardize Incident Response","slug":"how-sre-teams-use-runbooks-to-standardize-incident-response","published_at":"2026-07-11T10:05:17+00:00","page_url":"https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-runbooks-to-standardize-incident-response","show_page_url":"https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923","url":"https://audio.fexingo.com/business/the-site-reliability-podcast/episode-0104.mp3","audio_url":"https://audio.fexingo.com/business/the-site-reliability-podcast/episode-0104.mp3","summary":"In this episode of The Site Reliability Podcast, Lucas and Luna explore how SRE teams use runbooks to standardize incident response, reduce mean time to repair, and prevent cognitive overload during outages. They break down the anatomy of a good runbook—clear triggers, step-by-step diagnostic actions, escalation paths—and contrast it with the common failure mode of stale, never-updated documentation. Using real-world examples from large-scale production environments, they discuss how companies like Google and Netflix keep runbooks living documents through regular testing and automation. The hosts also touch on the tension between runbook rigidity and the need for human judgment in novel failures. If you've ever faced a pag er at 3 AM with no idea where to start, this episode is for you. #SiteReliabilityEngineering #SRE #IncidentResponse #Runbooks #DevOps #ProductionEngineering #Uptime #IncidentManagement #GoogleSRE #NetflixTech #OnCall #Automation #ITOperations #Technology #FexingoBusiness #BusinessPodcast #TechPodcast #ReliabilityEngineering Keep every episode free: buymeacoffee.com/fexingo","meta_description":"In this episode of The Site Reliability Podcast, Lucas and Luna explore how SRE teams use runbooks to standardize incident response, reduce mean time to r…","key_points":[],"chapters":[],"topics":[],"duration_seconds":721,"processing_state":"not_requested","actions":[{"name":"request_transcript","method":"POST","url":"https://stenobird.com/v1/public/podcasts/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/episodes/how-sre-teams-use-runbooks-to-standardize-incident-response/transcription-requests","description":"Idempotently request low-priority transcript generation for this episode."},{"name":"read_markdown","method":"GET","url":"https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-runbooks-to-standardize-incident-response.md","description":"Read the agent-friendly Markdown representation of this episode resource."}]}}