Episode

How SRE Teams Use Incident Metrics to Improve Postmortem Quality

Podcast
The Site Reliability Podcast with Fexingo: SRE, Uptime, and Production Engineering
Published
Jul 14, 2026
Duration seconds
585
Processing state
not_requested
Canonical source
https://audio.fexingo.com/business/the-site-reliability-podcast/episode-0110.mp3
Audio
https://audio.fexingo.com/business/the-site-reliability-podcast/episode-0110.mp3
JSON
/v1/public/podcasts/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/episodes/how-sre-teams-use-incident-metrics-to-improve-postmortem-quality
Markdown
/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-incident-metrics-to-improve-postmortem-quality.md

Actions

  • POST https://stenobird.com/v1/public/podcasts/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/episodes/how-sre-teams-use-incident-metrics-to-improve-postmortem-quality/transcription-requests
    Idempotently request low-priority transcript generation for this episode.
  • GET https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-incident-metrics-to-improve-postmortem-quality.md
    Read the agent-friendly Markdown representation of this episode resource.

Summary

Site reliability engineers write postmortems after every significant incident, but not all postmortems are created equal. In this episode, Lucas and Luna explore how leading SRE teams apply quantitative metrics — like time-to-detection, time-to-resolution, and mean time between incidents — to turn postmortems from reactive storytelling into a data-driven feedback loop. They examine a real case from a major cloud provider where metric-driven postmortems reduced repeat incidents by 40 percent over six quarters. Along the way, they discuss common pitfalls like survivorship bias in incident data and the surprising value of tracking the 'downtime cost per minute' to prioritize fixes. Whether your team writes five postmortems a month or five a quarter, you'll walk away with a concrete framework for measuring postmortem effectiveness. #SRE #SiteReliabilityEngineering #Postmortem #IncidentMetrics #MTTR #TimeToDetection #BlamelessCulture #Reliability #ProductionEngineering #Uptime #IncidentResponse #DataDriven #DevOps #Observability #CloudComputing #Technology #FexingoBusiness #BusinessPodcast Keep every episode free: buymeacoffee.com/fexingo