Episode
How SRE Teams Use Chaos Engineering to Find Hidden Failure Modes
- Published
- Jul 13, 2026
- Duration seconds
- 545
- Processing state
not_requested
Actions
POST https://stenobird.com/v1/public/podcasts/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/episodes/how-sre-teams-use-chaos-engineering-to-find-hidden-failure-modes/transcription-requests
Idempotently request low-priority transcript generation for this episode.GET https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-chaos-engineering-to-find-hidden-failure-modes.md
Read the agent-friendly Markdown representation of this episode resource.
Summary
In this episode of The Site Reliability Podcast, Lucas and Luna dive into chaos engineering as a proactive practice for uncovering hidden failure modes in production systems. They discuss how Netflix pioneered the approach with Chaos Monkey and how modern SRE teams use tools like Gremlin and Litmus to simulate failures — from CPU spikes to network partitions — in controlled experiments. The hosts explore the tension between running chaos experiments safely and the risk of unintended outages, sharing practical advice on starting small with low-blast-radius tests and using steady-state hypotheses to measure impact. They also touch on the cultural shift required: moving from a mindset of 'prevent every failure' to 'build systems that gracefully handle failure.' Listeners learn concrete steps to implement chaos engineering in their own teams. #ChaosEngineering #SiteReliabilityEngineering #SRE #FailureModeTesting #NetflixChaosMonkey #ProductionTesting #ResilienceEngineering #FaultInjection #Gremlin #LitmusChaos #SteadyStateHypothesis #BlastRadius #Technology #FexingoBusiness #BusinessPodcast #Uptime #IncidentPrevention #Observability Keep every episode free: buymeacoffee.com/fexingo