Episode
How Airbnb Uses Traffic Shifting to Prevent Cascading Failures
- Published
- Jul 10, 2026
- Duration seconds
- 497
- Processing state
not_requested
Actions
POST https://stenobird.com/v1/public/podcasts/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/episodes/how-airbnb-uses-traffic-shifting-to-prevent-cascading-failures/transcription-requests
Idempotently request low-priority transcript generation for this episode.GET https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-airbnb-uses-traffic-shifting-to-prevent-cascading-failures.md
Read the agent-friendly Markdown representation of this episode resource.
Summary
In this episode, Lucas and Luna explore how Airbnb's SRE team uses a technique called traffic shifting to prevent cascading failures during large-scale releases. They break down a real incident from 2024 where a problematic database migration was caught and rolled back in minutes thanks to gradual traffic redirection. Listeners learn the specific metrics Airbnb monitors—latency p99 and error rate—and how you can apply the same pattern in your own infrastructure. The hosts discuss the trade-offs of region-based vs. request-based shifting, and why this approach reduces blast radius without sacrificing deployment velocity. A practical deep dive for any engineer running production systems. #Airbnb #SiteReliabilityEngineering #TrafficShifting #CascadingFailures #IncidentResponse #DatabaseMigration #LatencyP99 #ErrorRate #BlastRadius #DeploymentVelocity #ProductionEngineering #SRE #Uptime #Technology #FexingoBusiness #BusinessPodcast #TheSiteReliabilityPodcast #TechOps Keep every episode free: buymeacoffee.com/fexingo