Episode
How SRE Teams Use Error Budgets to Balance Reliability and Velocity
- Published
- Jul 14, 2026
- Duration seconds
- 651
- Processing state
not_requested
Actions
POST https://stenobird.com/v1/public/podcasts/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/episodes/how-sre-teams-use-error-budgets-to-balance-reliability-and-velocity-3/transcription-requests
Idempotently request low-priority transcript generation for this episode.GET https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-error-budgets-to-balance-reliability-and-velocity-3.md
Read the agent-friendly Markdown representation of this episode resource.
Summary
Lucas and Luna dive into error budgets, the SRE mechanism that lets teams decide how much downtime is acceptable. They explore Google's original framework, how it resolves the tension between feature velocity and system reliability, and walk through a concrete example: a team with a 99.9% SLO gets 43 minutes of permissible downtime per month. They discuss what happens when the budget runs out, how to set SLOs that match user expectations, and why some teams still struggle to adopt error budgets despite the clarity they provide. No fluff—just the specifics of how a good error budget policy works in practice. #SRE #ErrorBudget #SiteReliabilityEngineering #SLO #GoogleSRE #ReliabilityEngineering #DevOps #IncidentManagement #FeatureVelocity #BlamelessCulture #ProductionEngineering #Technology #Uptime #TechOps #FexingoBusiness #BusinessPodcast #SiteReliability #TheSiteReliabilityPodcast Keep every episode free: buymeacoffee.com/fexingo