Episode
How SRE Teams Use Capacity Planning to Prevent Outages
- Published
- Jul 1, 2026
- Duration seconds
- 459
- Processing state
not_requested
Actions
POST https://stenobird.com/v1/public/podcasts/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/episodes/how-sre-teams-use-capacity-planning-to-prevent-outages/transcription-requests
Idempotently request low-priority transcript generation for this episode.GET https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-capacity-planning-to-prevent-outages.md
Read the agent-friendly Markdown representation of this episode resource.
Summary
In this episode of The Site Reliability Podcast, Lucas and Luna dive into capacity planning for SRE teams — the proactive discipline that keeps systems running when traffic spikes. Using the example of a major streaming platform's 2024 holiday season surge, they break down how capacity planning differs from simple scaling, why it's part of reliability engineering, and how teams use traffic forecasting, load testing, and cost models to avoid the 'one server too few' trap. They also discuss the tension between over-provisioning and under-provisioning, and how error budgets feed into capacity decisions. A concrete look at a less-glamorous but essential SRE practice. #CapacityPlanning #SiteReliabilityEngineering #SRE #Uptime #ProductionEngineering #Infrastructure #Scaling #LoadTesting #TrafficForecasting #ErrorBudget #CloudCost #Reliability #StreamingPlatform #HolidaySurge #TechPodcast #FexingoBusiness #BusinessPodcast #Technology Keep every episode free: buymeacoffee.com/fexingo