Episode
How SRE Teams Use Toil Budgets to Automate the Right Things
- Published
- Jul 16, 2026
- Duration seconds
- 616
- Processing state
not_requested
Actions
POST https://stenobird.com/v1/public/podcasts/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/episodes/how-sre-teams-use-toil-budgets-to-automate-the-right-things/transcription-requests
Idempotently request low-priority transcript generation for this episode.GET https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-toil-budgets-to-automate-the-right-things.md
Read the agent-friendly Markdown representation of this episode resource.
Summary
Episode 114 of The Site Reliability Podcast explores a concept that separates mature SRE teams from the rest: the toil budget. Lucas and Luna break down how teams at companies like Google and Shopify explicitly cap the amount of manual, repetitive work engineers spend each week — and what happens when you let teams decide which 30 percent of their toil to automate first. They walk through a real example from a mid-sized e-commerce platform that cut its on-call ticket volume by 60 percent after implementing a toil budget tied to a service-level objective. The episode also covers common failure modes: setting the budget too low, tracking toil wrong, and confusing cognitive overhead with toil. No vague advice — just a concrete framework any SRE team can steal. #SRE #SiteReliabilityEngineering #ToilBudget #Automation #GoogleSRE #DevOps #Uptime #IncidentResponse #OnCall #ServiceLevelObjective #PlatformEngineering #TechOps #ReliabilityEngineering #Observability #ProductionEngineering #IncidentManagement #FexingoBusiness #BusinessPodcast Keep every episode free: buymeacoffee.com/fexingo