{"podcast":{"title":"The Site Reliability Podcast with Fexingo: SRE, Uptime, and Production Engineering","slug":"the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923","podcast_index_feed_id":7871923,"rss_url":"https://feeds.fexingo.com/business/the-site-reliability-podcast.xml","website_url":"https://www.fexingo.com/","image_url":"https://audio.fexingo.com/business/the-site-reliability-podcast/cover.png","author":"Fexingo","episode_count":119,"summary":"Lucas and Luna cut through the noise around site reliability engineering to examine how real-world SRE teams balance uptime, incident response, and production change. Each episode takes a single concept — error budgets, toil automation, postmortem culture, capacity planning — and grounds it in a specific case: how a major streaming service reduced paging noise, how a payments platform rebuilt its incident command structure, or how a cloud provider manages multi-region failover. Lucas brings the numbers — latency percentiles, MTTR trends, SLO burn rates — while Luna pushes on the human and organizational trade-offs: What does a junior SRE need to know about on-call? How do you measure reliability without crushing innovation? Why do some blameless postmortems actually work? Together they treat SRE not as a certification topic but as a living practice, citing real outages, open-source tools, and engineering blogs. This show is for engineers, ops leads, and platform teams who already know the basics and want to debate the hard edges: Is 99.999% uptime always worth the cost? When should you deliberately degrade service to improve reliability? How do you design for resilience when your…","last_synced_at":"2026-07-18T22:18:41.635042+00:00","page_url":"https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923"},"episode":{"title":"How SRE Teams Use Game Days to Build Incident Muscle Memory","slug":"how-sre-teams-use-game-days-to-build-incident-muscle-memory","published_at":"2026-07-15T10:08:28+00:00","page_url":"https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-game-days-to-build-incident-muscle-memory","show_page_url":"https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923","url":"https://audio.fexingo.com/business/the-site-reliability-podcast/episode-0112.mp3","audio_url":"https://audio.fexingo.com/business/the-site-reliability-podcast/episode-0112.mp3","summary":"In this episode of The Site Reliability Podcast with Fexingo, Lucas and Luna dive into the practice of Game Days—simulated failure exercises that help SRE teams build muscle memory without risking production. They walk through how companies like Google and Netflix run these events, from chaos engineering experiments to tabletop exercises for on-call teams. The hosts discuss a specific case: a major e-commerce platform that discovered a critical database failover bug during a Game Day, preventing what would have been hours of downtime during Black Friday. They also cover practical tips for starting small, avoiding blame culture during simulations, and measuring success through mean time to detection and mean time to resolution improvements. If you're an SRE or platform engineer looking to harden your incident response, this episode offers concrete takeaways on designing effective Game Days that translate directly to real-world resilience. #GameDays #SRE #SiteReliabilityEngineering #ChaosEngineering #IncidentResponse #Uptime #ProductionEngineering #Resilience #FailureTesting #OnCall #MTTD #MTTR #Runbooks #GoogleSRE #Netflix #BlackFriday #Technology #FexingoBusiness Keep every episode free: buymeacoffee.com/fexingo","meta_description":"In this episode of The Site Reliability Podcast with Fexingo, Lucas and Luna dive into the practice of Game Days—simulated failure exercises that help SRE…","key_points":[],"chapters":[],"topics":[],"duration_seconds":500,"processing_state":"not_requested","actions":[{"name":"request_transcript","method":"POST","url":"https://stenobird.com/v1/public/podcasts/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/episodes/how-sre-teams-use-game-days-to-build-incident-muscle-memory/transcription-requests","description":"Idempotently request low-priority transcript generation for this episode."},{"name":"read_markdown","method":"GET","url":"https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-game-days-to-build-incident-muscle-memory.md","description":"Read the agent-friendly Markdown representation of this episode resource."}]}}