# How SRE Teams Use Chaos Engineering to Build Resilient Systems Page: https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-chaos-engineering-to-build-resilient-systems Text version: https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-chaos-engineering-to-build-resilient-systems.md Podcast: [The Site Reliability Podcast with Fexingo: SRE, Uptime, and Production Engineering](https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923) Published: 2026-06-30T22:31:33+00:00 Episode link: https://audio.fexingo.com/business/the-site-reliability-podcast/episode-0083.mp3 Audio file: https://audio.fexingo.com/business/the-site-reliability-podcast/episode-0083.mp3 Processing state: not_requested JSON: https://stenobird.com/v1/public/podcasts/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/episodes/how-sre-teams-use-chaos-engineering-to-build-resilient-systems Duration seconds: 705 ## Resource Lucas and Luna dive into chaos engineering, using Netflix's Chaos Monkey and the Simian Army as the prime example. Lucas explains how Netflix intentionally broke its own systems in production to uncover weaknesses before they caused real outages, citing the tool's origin story from 2011 and its evolution into a formal discipline. Luna challenges the notion that chaos experiments are too risky for smaller teams, and Lucas counters with the concept of a 'blast radius' and controlled experiments. They discuss how companies like Amazon and Google run game days that simulate infrastructure failures to verify redundancy and failover mechanisms. The episode walks through the principles of starting small, automating rollbacks, and measuring 'time to recover' as a key metric. The hosts close by reflecting on the broader cultural shift: chaos engineering is as much about team mindset as it is about tooling. #ChaosEngineering #SRE #Netflix #ChaosMonkey #SiteReliabilityEngineering #Resilience #DevOps #SiteReliability #FaultTolerance #ProductionEngineering #CloudComputing #Uptime #Infrastructure #Amazon #Google #Technology #FexingoBusiness #BusinessPodcast Keep every episode free: buymeacoffee.com/fexingo ## Actions - request_transcript: `POST https://stenobird.com/v1/public/podcasts/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/episodes/how-sre-teams-use-chaos-engineering-to-build-resilient-systems/transcription-requests` — Idempotently request low-priority transcript generation for this episode. - read_markdown: `GET https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-chaos-engineering-to-build-resilient-systems.md` — Read the agent-friendly Markdown representation of this episode resource. A page view does not enqueue transcription. Agents should invoke `request_transcript` explicitly when they need this episode processed. ## Transcript Full transcripts are not published on public pages unless there is a clear rights basis.