{"podcast":{"title":"The Site Reliability Podcast with Fexingo: SRE, Uptime, and Production Engineering","slug":"the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923","podcast_index_feed_id":7871923,"rss_url":"https://feeds.fexingo.com/business/the-site-reliability-podcast.xml","website_url":"https://www.fexingo.com/","image_url":"https://audio.fexingo.com/business/the-site-reliability-podcast/cover.png","author":"Fexingo","episode_count":119,"summary":"Lucas and Luna cut through the noise around site reliability engineering to examine how real-world SRE teams balance uptime, incident response, and production change. Each episode takes a single concept — error budgets, toil automation, postmortem culture, capacity planning — and grounds it in a specific case: how a major streaming service reduced paging noise, how a payments platform rebuilt its incident command structure, or how a cloud provider manages multi-region failover. Lucas brings the numbers — latency percentiles, MTTR trends, SLO burn rates — while Luna pushes on the human and organizational trade-offs: What does a junior SRE need to know about on-call? How do you measure reliability without crushing innovation? Why do some blameless postmortems actually work? Together they treat SRE not as a certification topic but as a living practice, citing real outages, open-source tools, and engineering blogs. This show is for engineers, ops leads, and platform teams who already know the basics and want to debate the hard edges: Is 99.999% uptime always worth the cost? When should you deliberately degrade service to improve reliability? How do you design for resilience when your…","last_synced_at":"2026-07-18T22:18:41.635042+00:00","page_url":"https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923"},"episode":{"title":"How SRE Teams Use Blameless Culture to Improve Incident Response","slug":"how-sre-teams-use-blameless-culture-to-improve-incident-response-2","published_at":"2026-07-07T22:28:24+00:00","page_url":"https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-blameless-culture-to-improve-incident-response-2","show_page_url":"https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923","url":"https://audio.fexingo.com/business/the-site-reliability-podcast/episode-0097.mp3","audio_url":"https://audio.fexingo.com/business/the-site-reliability-podcast/episode-0097.mp3","summary":"The most resilient systems aren't built by perfect engineers — they're built by teams that can learn from failure without fear of punishment. In this episode, Lucas and Luna explore the transition from 'root cause analysis' to 'blameless postmortems' at companies like Etsy and Google. They unpack how blameless culture actually works in practice: how Etsy's 'Blameless Postmortem' framework reduced repeat incidents by over 40 percent, why Google's SRE book explicitly bans the phrase 'human error', and how one incident at a major payment processor revealed deeper systemic issues that finger-pointing would have buried. The hosts also wrestle with the real tension — how do you hold engineers accountable without blame? And does 'blameless' ever become a free pass for carelessness? With concrete case studies and honest pushback, this episode gives SRE teams a practical guide to building a culture that actually learns. #BlamelessCulture #IncidentResponse #SRE #SiteReliabilityEngineering #Postmortem #LearningFromFailure #Etsy #Google #DevOps #ReliabilityEngineering #IncidentManagement #NoBlame #PsychologicalSafety #RootCauseAnalysis #ProductionEngineering #TechCulture #FexingoBusiness #BusinessPodcast Keep every episode free: buymeacoffee.com/fexingo","meta_description":"The most resilient systems aren't built by perfect engineers — they're built by teams that can learn from failure without fear of punishment. In this epis…","key_points":[],"chapters":[],"topics":[],"duration_seconds":651,"processing_state":"not_requested","actions":[{"name":"request_transcript","method":"POST","url":"https://stenobird.com/v1/public/podcasts/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/episodes/how-sre-teams-use-blameless-culture-to-improve-incident-response-2/transcription-requests","description":"Idempotently request low-priority transcript generation for this episode."},{"name":"read_markdown","method":"GET","url":"https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-blameless-culture-to-improve-incident-response-2.md","description":"Read the agent-friendly Markdown representation of this episode resource."}]}}