# How SRE Teams Use Incident Postmortems for Systemic Improvement Page: https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-incident-postmortems-for-systemic-improvement Text version: https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-incident-postmortems-for-systemic-improvement.md Podcast: [The Site Reliability Podcast with Fexingo: SRE, Uptime, and Production Engineering](https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923) Published: 2026-07-12T22:43:38+00:00 Episode link: https://audio.fexingo.com/business/the-site-reliability-podcast/episode-0107.mp3 Audio file: https://audio.fexingo.com/business/the-site-reliability-podcast/episode-0107.mp3 Processing state: not_requested JSON: https://stenobird.com/v1/public/podcasts/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/episodes/how-sre-teams-use-incident-postmortems-for-systemic-improvement Duration seconds: 634 ## Resource In this episode of The Site Reliability Podcast, Lucas and Luna explore how incident postmortems go beyond blame to drive systemic reliability improvements. They examine the anatomy of a good postmortem, including the 'five whys' technique, action item ownership, and the critical distinction between proximate cause and root cause. The hosts use the example of a major AWS outage in 2020—where a single misconfiguration cascaded across multiple services—to illustrate how even the best-run SRE teams can fall into the trap of superficial postmortems. They discuss how Etsy, Google, and other firms have turned postmortems into a core feedback loop for reliability engineering. The episode also covers common pitfalls like action item fatigue and the importance of blameless culture in fostering honest incident analysis. Listeners will learn practical steps to improve their own postmortem process, from timeline reconstruction to choosing the right level of detail. The hosts also touch on how automation and machine learning are beginning to augment human-led postmortem analysis, though they argue the human judgment element remains irreplaceable. The conversation closes with a reflection on how postmortems can shift a team's relationship with failure from fear to learning. #Postmortem #IncidentReview #SiteReliabilityEngineering #SRE #BlamelessCulture #RootCauseAnalysis #FiveWhys #AWSOutage #Etsy #GoogleSRE #ActionItem #LearningFromFailure #ResilienceEngineering #Technology #FexingoBusiness #BusinessPodcast #IncidentManagement #Reliability Keep every episode free: buymeacoffee.com/fexingo ## Actions - request_transcript: `POST https://stenobird.com/v1/public/podcasts/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/episodes/how-sre-teams-use-incident-postmortems-for-systemic-improvement/transcription-requests` — Idempotently request low-priority transcript generation for this episode. - read_markdown: `GET https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-incident-postmortems-for-systemic-improvement.md` — Read the agent-friendly Markdown representation of this episode resource. A page view does not enqueue transcription. Agents should invoke `request_transcript` explicitly when they need this episode processed. ## Transcript Full transcripts are not published on public pages unless there is a clear rights basis.