# How SRE Teams Use Fault Tree Analysis to Prevent Root Causes Page: https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-fault-tree-analysis-to-prevent-root-causes Text version: https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-fault-tree-analysis-to-prevent-root-causes.md Podcast: [The Site Reliability Podcast with Fexingo: SRE, Uptime, and Production Engineering](https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923) Published: 2026-06-25T09:58:05+00:00 Episode link: https://audio.fexingo.com/business/the-site-reliability-podcast/episode-0072.mp3 Audio file: https://audio.fexingo.com/business/the-site-reliability-podcast/episode-0072.mp3 Processing state: not_requested JSON: https://stenobird.com/v1/public/podcasts/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/episodes/how-sre-teams-use-fault-tree-analysis-to-prevent-root-causes Duration seconds: 709 ## Resource In this episode of The Site Reliability Podcast, Lucas and Luna explore how SRE teams apply fault tree analysis (FTA) from aerospace and nuclear engineering to reduce incident recurrence. Using a real 2025 outage at a major streaming platform where a cascading DNS failure took down services for 47 minutes, they break down the top-down logic of FTA, how it differs from postmortem 5 whys, and why teams at companies like Netflix and Amazon are adopting it alongside their existing incident analysis toolkit. Lucas explains the basic mechanics: starting with an undesired top event, then decomposing it into intermediate and basic events united by AND/OR gates, calculating probability cut sets, and using the results to prioritize mitigations. Luna pushes back on the learning curve and time investment, and they discuss lightweight variants like 'bow-tie analysis' used by smaller teams. The episode also touches on common pitfalls: confirmation bias, data quality, and treating a first-pass FTA as finished. Practical takeaway: any team can run a two-hour FTA workshop after a major incident to surface hidden single points of failure. #FaultTreeAnalysis #SRE #SiteReliabilityEngineering #RootCauseAnalysis #IncidentManagement #ReliabilityEngineering #Postmortem #5Whys #BowTieAnalysis #DNSOutage #CascadingFailure #Netflix #Amazon #ProductionEngineering #Technology #FexingoBusiness #BusinessPodcast #SREPodcast Keep every episode free: buymeacoffee.com/fexingo ## Actions - request_transcript: `POST https://stenobird.com/v1/public/podcasts/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/episodes/how-sre-teams-use-fault-tree-analysis-to-prevent-root-causes/transcription-requests` — Idempotently request low-priority transcript generation for this episode. - read_markdown: `GET https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-fault-tree-analysis-to-prevent-root-causes.md` — Read the agent-friendly Markdown representation of this episode resource. A page view does not enqueue transcription. Agents should invoke `request_transcript` explicitly when they need this episode processed. ## Transcript Full transcripts are not published on public pages unless there is a clear rights basis.