{"podcast":{"title":"The Site Reliability Podcast with Fexingo: SRE, Uptime, and Production Engineering","slug":"the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923","podcast_index_feed_id":7871923,"rss_url":"https://feeds.fexingo.com/business/the-site-reliability-podcast.xml","website_url":"https://www.fexingo.com/","image_url":"https://audio.fexingo.com/business/the-site-reliability-podcast/cover.png","author":"Fexingo","episode_count":119,"summary":"Lucas and Luna cut through the noise around site reliability engineering to examine how real-world SRE teams balance uptime, incident response, and production change. Each episode takes a single concept — error budgets, toil automation, postmortem culture, capacity planning — and grounds it in a specific case: how a major streaming service reduced paging noise, how a payments platform rebuilt its incident command structure, or how a cloud provider manages multi-region failover. Lucas brings the numbers — latency percentiles, MTTR trends, SLO burn rates — while Luna pushes on the human and organizational trade-offs: What does a junior SRE need to know about on-call? How do you measure reliability without crushing innovation? Why do some blameless postmortems actually work? Together they treat SRE not as a certification topic but as a living practice, citing real outages, open-source tools, and engineering blogs. This show is for engineers, ops leads, and platform teams who already know the basics and want to debate the hard edges: Is 99.999% uptime always worth the cost? When should you deliberately degrade service to improve reliability? How do you design for resilience when your…","last_synced_at":"2026-07-18T22:18:41.635042+00:00","page_url":"https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923"},"episode":{"title":"How SRE Teams Use Load Shedding to Protect Critical Services","slug":"how-sre-teams-use-load-shedding-to-protect-critical-services","published_at":"2026-07-16T22:32:03+00:00","page_url":"https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-load-shedding-to-protect-critical-services","show_page_url":"https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923","url":"https://audio.fexingo.com/business/the-site-reliability-podcast/episode-0115.mp3","audio_url":"https://audio.fexingo.com/business/the-site-reliability-podcast/episode-0115.mp3","summary":"Site reliability engineers have a last-resort play when traffic overwhelms their systems: load shedding. In this episode, Lucas and Luna explore how companies like Google and Amazon use intentional, tiered request-dropping to keep their most critical services alive during demand spikes. They break down the difference between load shedding and rate limiting, discuss why many teams get the prioritization wrong, and walk through a real-world example involving a streaming platform's checkout service during a major sale event. Listeners will learn how to design priority-aware request queues, why every microservice should have a defined shedding order, and how to test load shedding in game days without causing real downtime. This episode offers a concrete playbook for any SRE team looking to answer the question: when you have to say no, whom do you say no to first? #LoadShedding #SiteReliabilityEngineering #SRE #ProductionEngineering #TrafficManagement #PriorityQueues #CapacityPlanning #ResilienceEngineering #IncidentResponse #FaultTolerance #GoogleSRE #AmazonPrimeDay #CheckoutService #GameDays #Uptime #TechOps #Technology #FexingoBusiness Keep every episode free: buymeacoffee.com/fexingo","meta_description":"Site reliability engineers have a last-resort play when traffic overwhelms their systems: load shedding. In this episode, Lucas and Luna explore how compa…","key_points":[],"chapters":[],"topics":[],"duration_seconds":753,"processing_state":"not_requested","actions":[{"name":"request_transcript","method":"POST","url":"https://stenobird.com/v1/public/podcasts/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/episodes/how-sre-teams-use-load-shedding-to-protect-critical-services/transcription-requests","description":"Idempotently request low-priority transcript generation for this episode."},{"name":"read_markdown","method":"GET","url":"https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-load-shedding-to-protect-critical-services.md","description":"Read the agent-friendly Markdown representation of this episode resource."}]}}