{"podcast":{"title":"The Site Reliability Podcast with Fexingo: SRE, Uptime, and Production Engineering","slug":"the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923","podcast_index_feed_id":7871923,"rss_url":"https://feeds.fexingo.com/business/the-site-reliability-podcast.xml","website_url":"https://www.fexingo.com/","image_url":"https://audio.fexingo.com/business/the-site-reliability-podcast/cover.png","author":"Fexingo","episode_count":119,"summary":"Lucas and Luna cut through the noise around site reliability engineering to examine how real-world SRE teams balance uptime, incident response, and production change. Each episode takes a single concept — error budgets, toil automation, postmortem culture, capacity planning — and grounds it in a specific case: how a major streaming service reduced paging noise, how a payments platform rebuilt its incident command structure, or how a cloud provider manages multi-region failover. Lucas brings the numbers — latency percentiles, MTTR trends, SLO burn rates — while Luna pushes on the human and organizational trade-offs: What does a junior SRE need to know about on-call? How do you measure reliability without crushing innovation? Why do some blameless postmortems actually work? Together they treat SRE not as a certification topic but as a living practice, citing real outages, open-source tools, and engineering blogs. This show is for engineers, ops leads, and platform teams who already know the basics and want to debate the hard edges: Is 99.999% uptime always worth the cost? When should you deliberately degrade service to improve reliability? How do you design for resilience when your…","last_synced_at":"2026-07-18T22:18:41.635042+00:00","page_url":"https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923"},"episode":{"title":"How SRE Teams Use Chaos Engineering to Build Resilient Systems","slug":"how-sre-teams-use-chaos-engineering-to-build-resilient-systems","published_at":"2026-06-30T22:31:33+00:00","page_url":"https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-chaos-engineering-to-build-resilient-systems","show_page_url":"https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923","url":"https://audio.fexingo.com/business/the-site-reliability-podcast/episode-0083.mp3","audio_url":"https://audio.fexingo.com/business/the-site-reliability-podcast/episode-0083.mp3","summary":"Lucas and Luna dive into chaos engineering, using Netflix's Chaos Monkey and the Simian Army as the prime example. Lucas explains how Netflix intentionally broke its own systems in production to uncover weaknesses before they caused real outages, citing the tool's origin story from 2011 and its evolution into a formal discipline. Luna challenges the notion that chaos experiments are too risky for smaller teams, and Lucas counters with the concept of a 'blast radius' and controlled experiments. They discuss how companies like Amazon and Google run game days that simulate infrastructure failures to verify redundancy and failover mechanisms. The episode walks through the principles of starting small, automating rollbacks, and measuring 'time to recover' as a key metric. The hosts close by reflecting on the broader cultural shift: chaos engineering is as much about team mindset as it is about tooling. #ChaosEngineering #SRE #Netflix #ChaosMonkey #SiteReliabilityEngineering #Resilience #DevOps #SiteReliability #FaultTolerance #ProductionEngineering #CloudComputing #Uptime #Infrastructure #Amazon #Google #Technology #FexingoBusiness #BusinessPodcast Keep every episode free: buymeacoffee.com/fexingo","meta_description":"Lucas and Luna dive into chaos engineering, using Netflix's Chaos Monkey and the Simian Army as the prime example. Lucas explains how Netflix intentionall…","key_points":[],"chapters":[],"topics":[],"duration_seconds":705,"processing_state":"not_requested","actions":[{"name":"request_transcript","method":"POST","url":"https://stenobird.com/v1/public/podcasts/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/episodes/how-sre-teams-use-chaos-engineering-to-build-resilient-systems/transcription-requests","description":"Idempotently request low-priority transcript generation for this episode."},{"name":"read_markdown","method":"GET","url":"https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-chaos-engineering-to-build-resilient-systems.md","description":"Read the agent-friendly Markdown representation of this episode resource."}]}}