Episode

How SRE Teams Use Dependency Graphs to Prevent Cascading Failures

Podcast
The Site Reliability Podcast with Fexingo: SRE, Uptime, and Production Engineering
Published
Jul 9, 2026
Duration seconds
605
Processing state
not_requested
Canonical source
https://audio.fexingo.com/business/the-site-reliability-podcast/episode-0101.mp3
Audio
https://audio.fexingo.com/business/the-site-reliability-podcast/episode-0101.mp3
JSON
/v1/public/podcasts/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/episodes/how-sre-teams-use-dependency-graphs-to-prevent-cascading-failures
Markdown
/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-dependency-graphs-to-prevent-cascading-failures.md

Actions

  • POST https://stenobird.com/v1/public/podcasts/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/episodes/how-sre-teams-use-dependency-graphs-to-prevent-cascading-failures/transcription-requests
    Idempotently request low-priority transcript generation for this episode.
  • GET https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-dependency-graphs-to-prevent-cascading-failures.md
    Read the agent-friendly Markdown representation of this episode resource.

Summary

In this episode of The Site Reliability Podcast, Lucas and Luna explore how SRE teams build and use dependency graphs to map service interconnections and prevent cascading failures. They discuss the 2021 Fastly outage as a real-world example of a dependency chain gone wrong, and walk through how teams at companies like Google and Netflix maintain dynamic dependency maps using tools like service meshes and distributed tracing. Lucas explains the difference between static and runtime dependency graphs, and Luna shares how a dependency graph helped her team catch a looming database overload during a Black Friday simulation. The hosts also touch on the challenges of graph maintenance at scale and the emerging role of AI in automating dependency discovery. A practical, concrete look at how understanding your system's hidden connections can save your service. #SiteReliabilityEngineering #SRE #DependencyGraphs #CascadingFailures #Fastly #ServiceMesh #DistributedTracing #ResilienceEngineering #ProductionEngineering #TechPodcast #IncidentResponse #SystemDesign #Observability #Netflix #Google #Uptime #FexingoBusiness #BusinessPodcast Keep every episode free: buymeacoffee.com/fexingo