{"podcast":{"title":"The Site Reliability Podcast with Fexingo: SRE, Uptime, and Production Engineering","slug":"the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923","podcast_index_feed_id":7871923,"rss_url":"https://feeds.fexingo.com/business/the-site-reliability-podcast.xml","website_url":"https://www.fexingo.com/","image_url":"https://audio.fexingo.com/business/the-site-reliability-podcast/cover.png","author":"Fexingo","episode_count":119,"summary":"Lucas and Luna cut through the noise around site reliability engineering to examine how real-world SRE teams balance uptime, incident response, and production change. Each episode takes a single concept — error budgets, toil automation, postmortem culture, capacity planning — and grounds it in a specific case: how a major streaming service reduced paging noise, how a payments platform rebuilt its incident command structure, or how a cloud provider manages multi-region failover. Lucas brings the numbers — latency percentiles, MTTR trends, SLO burn rates — while Luna pushes on the human and organizational trade-offs: What does a junior SRE need to know about on-call? How do you measure reliability without crushing innovation? Why do some blameless postmortems actually work? Together they treat SRE not as a certification topic but as a living practice, citing real outages, open-source tools, and engineering blogs. This show is for engineers, ops leads, and platform teams who already know the basics and want to debate the hard edges: Is 99.999% uptime always worth the cost? When should you deliberately degrade service to improve reliability? How do you design for resilience when your…","last_synced_at":"2026-07-18T22:18:41.635042+00:00","page_url":"https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923"},"episode":{"title":"How SRE Teams Use Error Budgets to Balance Reliability and Velocity","slug":"how-sre-teams-use-error-budgets-to-balance-reliability-and-velocity","published_at":"2026-07-04T10:27:23+00:00","page_url":"https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-error-budgets-to-balance-reliability-and-velocity","show_page_url":"https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923","url":"https://audio.fexingo.com/business/the-site-reliability-podcast/episode-0090.mp3","audio_url":"https://audio.fexingo.com/business/the-site-reliability-podcast/episode-0090.mp3","summary":"In this episode, Lucas and Luna dive into the concept of error budgets—a cornerstone of Site Reliability Engineering that defines how much unreliability a team can tolerate while still meeting their Service Level Objectives. They explore how error budgets help SRE teams make data-driven trade-offs between shipping new features and maintaining system stability. Using examples from Google's original SRE model and real-world applications at companies like Netflix and Etsy, they unpack how tracking error budget burn rates can trigger automated rollbacks or throttle deployments. Lucas breaks down the math behind error budgets, explaining how they derive from SLOs and how teams calculate budget consumption over time. The conversation also covers common pitfalls, like teams setting error budgets too tight or ignoring the budget entirely during crunch time. By the end, listeners will understand why error budgets are not just a monitoring tool but a cultural mechanism that aligns engineering incentives with business priorities. Tune in to learn how to use error budgets to ship faster with confidence on The Site Reliability Podcast with Fexingo. #ErrorBudget #SRE #SiteReliabilityEngineering #ServiceLevelObjective #SLO #Reliability #Velocity #IncidentResponse #GoogleSRE #Netflix #Etsy #DeploymentAutomation #ToilBudget #EngineeringCulture #TechPodcast #Technology #FexingoBusiness #BusinessPodcast Keep every episode free: buymeacoffee.com/fexingo","meta_description":"In this episode, Lucas and Luna dive into the concept of error budgets—a cornerstone of Site Reliability Engineering that defines how much unreliability a…","key_points":[],"chapters":[],"topics":[],"duration_seconds":709,"processing_state":"not_requested","actions":[{"name":"request_transcript","method":"POST","url":"https://stenobird.com/v1/public/podcasts/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/episodes/how-sre-teams-use-error-budgets-to-balance-reliability-and-velocity/transcription-requests","description":"Idempotently request low-priority transcript generation for this episode."},{"name":"read_markdown","method":"GET","url":"https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-error-budgets-to-balance-reliability-and-velocity.md","description":"Read the agent-friendly Markdown representation of this episode resource."}]}}