{"podcast":{"title":"The Site Reliability Podcast with Fexingo: SRE, Uptime, and Production Engineering","slug":"the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923","podcast_index_feed_id":7871923,"rss_url":"https://feeds.fexingo.com/business/the-site-reliability-podcast.xml","website_url":"https://www.fexingo.com/","image_url":"https://audio.fexingo.com/business/the-site-reliability-podcast/cover.png","author":"Fexingo","episode_count":119,"summary":"Lucas and Luna cut through the noise around site reliability engineering to examine how real-world SRE teams balance uptime, incident response, and production change. Each episode takes a single concept — error budgets, toil automation, postmortem culture, capacity planning — and grounds it in a specific case: how a major streaming service reduced paging noise, how a payments platform rebuilt its incident command structure, or how a cloud provider manages multi-region failover. Lucas brings the numbers — latency percentiles, MTTR trends, SLO burn rates — while Luna pushes on the human and organizational trade-offs: What does a junior SRE need to know about on-call? How do you measure reliability without crushing innovation? Why do some blameless postmortems actually work? Together they treat SRE not as a certification topic but as a living practice, citing real outages, open-source tools, and engineering blogs. This show is for engineers, ops leads, and platform teams who already know the basics and want to debate the hard edges: Is 99.999% uptime always worth the cost? When should you deliberately degrade service to improve reliability? How do you design for resilience when your…","last_synced_at":"2026-07-18T22:18:41.635042+00:00","page_url":"https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923"},"episode":{"title":"How SRE Teams Use Cost of Delay to Prioritize Reliability Work","slug":"how-sre-teams-use-cost-of-delay-to-prioritize-reliability-work","published_at":"2026-06-30T10:18:15+00:00","page_url":"https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-cost-of-delay-to-prioritize-reliability-work","show_page_url":"https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923","url":"https://audio.fexingo.com/business/the-site-reliability-podcast/episode-0082.mp3","audio_url":"https://audio.fexingo.com/business/the-site-reliability-podcast/episode-0082.mp3","summary":"Episode 82 of The Site Reliability Podcast examines how cost of delay — a concept borrowed from product development — helps SRE teams decide which reliability projects to tackle first. Lucas and Luna walk through a real example from a mid-sized fintech company that used cost of delay to justify migrating from a legacy database to a distributed SQL solution. The episode explains how to calculate cost of delay per unit of time, how to weigh it against implementation effort, and why this framework prevents teams from wasting weeks on marginal improvements while ignoring critical risks. The hosts also discuss common pitfalls, like treating all outages as equal in cost, and how to adjust the model for incidents that erode trust rather than just revenue. By the end, listeners will have a concrete method they can adapt for their own SLO-based prioritization decisions. #CostOfDelay #SRE #SiteReliabilityEngineering #Prioritization #Reliability #IncidentResponse #SLO #ServiceLevelObjectives #Fintech #DatabaseMigration #DistributedSQL #RiskManagement #ProductManagement #TechOps #ProductionEngineering #FexingoBusiness #BusinessPodcast #Technology Keep every episode free: buymeacoffee.com/fexingo","meta_description":"Episode 82 of The Site Reliability Podcast examines how cost of delay — a concept borrowed from product development — helps SRE teams decide which reliabi…","key_points":[],"chapters":[],"topics":[],"duration_seconds":768,"processing_state":"not_requested","actions":[{"name":"request_transcript","method":"POST","url":"https://stenobird.com/v1/public/podcasts/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/episodes/how-sre-teams-use-cost-of-delay-to-prioritize-reliability-work/transcription-requests","description":"Idempotently request low-priority transcript generation for this episode."},{"name":"read_markdown","method":"GET","url":"https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-cost-of-delay-to-prioritize-reliability-work.md","description":"Read the agent-friendly Markdown representation of this episode resource."}]}}