Episode

How SRE Teams Use Cost of Delay to Prioritize Reliability Work

Podcast
The Site Reliability Podcast with Fexingo: SRE, Uptime, and Production Engineering
Published
Jun 30, 2026
Duration seconds
768
Processing state
not_requested
Canonical source
https://audio.fexingo.com/business/the-site-reliability-podcast/episode-0082.mp3
Audio
https://audio.fexingo.com/business/the-site-reliability-podcast/episode-0082.mp3
JSON
/v1/public/podcasts/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/episodes/how-sre-teams-use-cost-of-delay-to-prioritize-reliability-work
Markdown
/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-cost-of-delay-to-prioritize-reliability-work.md

Actions

  • POST https://stenobird.com/v1/public/podcasts/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/episodes/how-sre-teams-use-cost-of-delay-to-prioritize-reliability-work/transcription-requests
    Idempotently request low-priority transcript generation for this episode.
  • GET https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-cost-of-delay-to-prioritize-reliability-work.md
    Read the agent-friendly Markdown representation of this episode resource.

Summary

Episode 82 of The Site Reliability Podcast examines how cost of delay — a concept borrowed from product development — helps SRE teams decide which reliability projects to tackle first. Lucas and Luna walk through a real example from a mid-sized fintech company that used cost of delay to justify migrating from a legacy database to a distributed SQL solution. The episode explains how to calculate cost of delay per unit of time, how to weigh it against implementation effort, and why this framework prevents teams from wasting weeks on marginal improvements while ignoring critical risks. The hosts also discuss common pitfalls, like treating all outages as equal in cost, and how to adjust the model for incidents that erode trust rather than just revenue. By the end, listeners will have a concrete method they can adapt for their own SLO-based prioritization decisions. #CostOfDelay #SRE #SiteReliabilityEngineering #Prioritization #Reliability #IncidentResponse #SLO #ServiceLevelObjectives #Fintech #DatabaseMigration #DistributedSQL #RiskManagement #ProductManagement #TechOps #ProductionEngineering #FexingoBusiness #BusinessPodcast #Technology Keep every episode free: buymeacoffee.com/fexingo