{"podcast":{"title":"The Site Reliability Podcast with Fexingo: SRE, Uptime, and Production Engineering","slug":"the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923","podcast_index_feed_id":7871923,"rss_url":"https://feeds.fexingo.com/business/the-site-reliability-podcast.xml","website_url":"https://www.fexingo.com/","image_url":"https://audio.fexingo.com/business/the-site-reliability-podcast/cover.png","author":"Fexingo","episode_count":119,"summary":"Lucas and Luna cut through the noise around site reliability engineering to examine how real-world SRE teams balance uptime, incident response, and production change. Each episode takes a single concept — error budgets, toil automation, postmortem culture, capacity planning — and grounds it in a specific case: how a major streaming service reduced paging noise, how a payments platform rebuilt its incident command structure, or how a cloud provider manages multi-region failover. Lucas brings the numbers — latency percentiles, MTTR trends, SLO burn rates — while Luna pushes on the human and organizational trade-offs: What does a junior SRE need to know about on-call? How do you measure reliability without crushing innovation? Why do some blameless postmortems actually work? Together they treat SRE not as a certification topic but as a living practice, citing real outages, open-source tools, and engineering blogs. This show is for engineers, ops leads, and platform teams who already know the basics and want to debate the hard edges: Is 99.999% uptime always worth the cost? When should you deliberately degrade service to improve reliability? How do you design for resilience when your…","last_synced_at":"2026-07-18T22:18:41.635042+00:00","page_url":"https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923"},"episode":{"title":"How SRE Teams Use Fault Tree Analysis to Prevent Root Causes","slug":"how-sre-teams-use-fault-tree-analysis-to-prevent-root-causes","published_at":"2026-06-25T09:58:05+00:00","page_url":"https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-fault-tree-analysis-to-prevent-root-causes","show_page_url":"https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923","url":"https://audio.fexingo.com/business/the-site-reliability-podcast/episode-0072.mp3","audio_url":"https://audio.fexingo.com/business/the-site-reliability-podcast/episode-0072.mp3","summary":"In this episode of The Site Reliability Podcast, Lucas and Luna explore how SRE teams apply fault tree analysis (FTA) from aerospace and nuclear engineering to reduce incident recurrence. Using a real 2025 outage at a major streaming platform where a cascading DNS failure took down services for 47 minutes, they break down the top-down logic of FTA, how it differs from postmortem 5 whys, and why teams at companies like Netflix and Amazon are adopting it alongside their existing incident analysis toolkit. Lucas explains the basic mechanics: starting with an undesired top event, then decomposing it into intermediate and basic events united by AND/OR gates, calculating probability cut sets, and using the results to prioritize mitigations. Luna pushes back on the learning curve and time investment, and they discuss lightweight variants like 'bow-tie analysis' used by smaller teams. The episode also touches on common pitfalls: confirmation bias, data quality, and treating a first-pass FTA as finished. Practical takeaway: any team can run a two-hour FTA workshop after a major incident to surface hidden single points of failure. #FaultTreeAnalysis #SRE #SiteReliabilityEngineering #RootCauseAnalysis #IncidentManagement #ReliabilityEngineering #Postmortem #5Whys #BowTieAnalysis #DNSOutage #CascadingFailure #Netflix #Amazon #ProductionEngineering #Technology #FexingoBusiness #BusinessPodcast #SREPodcast Keep every episode free: buymeacoffee.com/fexingo","meta_description":"In this episode of The Site Reliability Podcast, Lucas and Luna explore how SRE teams apply fault tree analysis (FTA) from aerospace and nuclear engineeri…","key_points":[],"chapters":[],"topics":[],"duration_seconds":709,"processing_state":"not_requested","actions":[{"name":"request_transcript","method":"POST","url":"https://stenobird.com/v1/public/podcasts/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/episodes/how-sre-teams-use-fault-tree-analysis-to-prevent-root-causes/transcription-requests","description":"Idempotently request low-priority transcript generation for this episode."},{"name":"read_markdown","method":"GET","url":"https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-fault-tree-analysis-to-prevent-root-causes.md","description":"Read the agent-friendly Markdown representation of this episode resource."}]}}