Episode
How SRE Teams Use Readiness Checks to Prevent Bad Deployments
- Published
- Jun 22, 2026
- Duration seconds
- 487
- Processing state
not_requested
Actions
POST https://stenobird.com/v1/public/podcasts/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/episodes/how-sre-teams-use-readiness-checks-to-prevent-bad-deployments/transcription-requests
Idempotently request low-priority transcript generation for this episode.GET https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-readiness-checks-to-prevent-bad-deployments.md
Read the agent-friendly Markdown representation of this episode resource.
Summary
Site reliability teams spend huge effort on monitoring and alerting—but some of the worst outages start the moment a deployment goes live. In this episode, Lucas and Luna break down how readiness checks, or health probes, act as the first line of defense against bad code reaching production. Using the example of a major Kubernetes rollout gone wrong at a large e-commerce company, they explain the difference between liveness and readiness probes, when to use startup probes, and why gating deployments on real traffic behavior—not just process completion—can prevent cascading failures. They also discuss common pitfalls like treating readiness checks as testing substitutes and conflating service health with instance health. If you've ever wondered why a seemingly simple deploy triggered a five-alarm incident, this episode is for you. #SRE #SiteReliabilityEngineering #Kubernetes #ReadinessProbes #LivenessProbes #DeploymentSafety #HealthChecks #IncidentPrevention #CloudNative #ProductionEngineering #DevOps #CI/CD #DeploymentGates #CanaryDeployments #Technology #FexingoBusiness #BusinessPodcast #TheSiteReliabilityPodcast Keep every episode free: buymeacoffee.com/fexingo