Episode
Into the VOID Report With Casey Rosenthal and Courtney Nash
- Podcast
- Arrested DevOps
- Published
- Apr 13, 2023
- Duration seconds
- 3590
- Processing state
processed- Canonical source
- https://www.arresteddevops.com/into-the-void/
Actions
POST https://stenobird.com/v1/public/podcasts/arrested-devops/episodes/into-the-void-report-with-casey-rosenthal-and-courtney-nash/transcription-requests
Idempotently request low-priority transcript generation for this episode.GET https://stenobird.com/podcast/arrested-devops/into-the-void-report-with-casey-rosenthal-and-courtney-nash.md
Read the agent-friendly Markdown representation of this episode resource.
Summary
An analysis of the Verica Open Incident Database (VOID) report, exploring how modern software complexity challenges traditional incident response. The discussion critiques the use of bureaucratic metrics like MTTR and examines the shift from Root Cause Analysis to learning-oriented post-incident reviews.
Topics
- DevOps
- Incident Response
- Resilience Engineering
- Software Reliability
- Observability
- Site Reliability Engineering
- Post-Mortem Analysis
- IT Governance
Highlights
- Main idea: The VOID report highlights the need to commoditize incident knowledge to improve industry-wide resilience
- Failure mode: Using metrics tied to compensation, such as MTTR, incentivizes teams to litigate severity rather than focus on system safety
- Practical takeaway: Organizations should transition from 'Root Cause Analysis' to 'Post-Incident Reviews' to prioritize depth and learning over blame
- Main idea: Software engineering has inadvertently become a 'bureaucratic profession,' adopting manufacturing-style hierarchies that hinder agility
- Failure mode: Defining incident severity based solely on duration or revenue loss fails to account for the complexities of distributed, global-scale systems
Chapters
5:20The VOID Report Origins: An introduction to the Verica Open Incident Database and its mission to share incident data openly.9:50Sharing Incident Knowledge: Discussing the cultural shift needed to move away from adversarial information silos toward shared learning.19:00Evolution of DevOps Practices: Reflecting on how practices have evolved from simple automated testing to managing complex, distributed software environments.27:50The Perils of Metric-Driven IT: How treating IT as a cost center leads to the use of flawed metrics like MTTR that undermine true reliability.32:20Software as a Bureaucratic Profession: A critique of how modern software organizations have adopted rigid, bureaucratic structures that conflict with system safety.45:40Measuring Incident Impact: Examining how incident reports are used for training and code reviews rather than just tracking failures.50:05The Shift to Post-Incident Reviews: Analyzing the move away from RCA (Root Cause Analysis) toward more qualitative, deep-dive incident reviews.