Episode
How SRE Teams Use Dependency Graphs to Prevent Cascading Failures
- Published
- Jul 9, 2026
- Duration seconds
- 605
- Processing state
not_requested
Actions
POST https://stenobird.com/v1/public/podcasts/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/episodes/how-sre-teams-use-dependency-graphs-to-prevent-cascading-failures/transcription-requests
Idempotently request low-priority transcript generation for this episode.GET https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-dependency-graphs-to-prevent-cascading-failures.md
Read the agent-friendly Markdown representation of this episode resource.
Summary
In this episode of The Site Reliability Podcast, Lucas and Luna explore how SRE teams build and use dependency graphs to map service interconnections and prevent cascading failures. They discuss the 2021 Fastly outage as a real-world example of a dependency chain gone wrong, and walk through how teams at companies like Google and Netflix maintain dynamic dependency maps using tools like service meshes and distributed tracing. Lucas explains the difference between static and runtime dependency graphs, and Luna shares how a dependency graph helped her team catch a looming database overload during a Black Friday simulation. The hosts also touch on the challenges of graph maintenance at scale and the emerging role of AI in automating dependency discovery. A practical, concrete look at how understanding your system's hidden connections can save your service. #SiteReliabilityEngineering #SRE #DependencyGraphs #CascadingFailures #Fastly #ServiceMesh #DistributedTracing #ResilienceEngineering #ProductionEngineering #TechPodcast #IncidentResponse #SystemDesign #Observability #Netflix #Google #Uptime #FexingoBusiness #BusinessPodcast Keep every episode free: buymeacoffee.com/fexingo