Episode

How SRE Teams Use Readiness Checks to Prevent Bad Deployments

Podcast
The Site Reliability Podcast with Fexingo: SRE, Uptime, and Production Engineering
Published
Jun 22, 2026
Duration seconds
487
Processing state
not_requested
Canonical source
https://audio.fexingo.com/business/the-site-reliability-podcast/episode-0066.mp3
Audio
https://audio.fexingo.com/business/the-site-reliability-podcast/episode-0066.mp3
JSON
/v1/public/podcasts/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/episodes/how-sre-teams-use-readiness-checks-to-prevent-bad-deployments
Markdown
/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-readiness-checks-to-prevent-bad-deployments.md

Actions

  • POST https://stenobird.com/v1/public/podcasts/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/episodes/how-sre-teams-use-readiness-checks-to-prevent-bad-deployments/transcription-requests
    Idempotently request low-priority transcript generation for this episode.
  • GET https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-readiness-checks-to-prevent-bad-deployments.md
    Read the agent-friendly Markdown representation of this episode resource.

Summary

Site reliability teams spend huge effort on monitoring and alerting—but some of the worst outages start the moment a deployment goes live. In this episode, Lucas and Luna break down how readiness checks, or health probes, act as the first line of defense against bad code reaching production. Using the example of a major Kubernetes rollout gone wrong at a large e-commerce company, they explain the difference between liveness and readiness probes, when to use startup probes, and why gating deployments on real traffic behavior—not just process completion—can prevent cascading failures. They also discuss common pitfalls like treating readiness checks as testing substitutes and conflating service health with instance health. If you've ever wondered why a seemingly simple deploy triggered a five-alarm incident, this episode is for you. #SRE #SiteReliabilityEngineering #Kubernetes #ReadinessProbes #LivenessProbes #DeploymentSafety #HealthChecks #IncidentPrevention #CloudNative #ProductionEngineering #DevOps #CI/CD #DeploymentGates #CanaryDeployments #Technology #FexingoBusiness #BusinessPodcast #TheSiteReliabilityPodcast Keep every episode free: buymeacoffee.com/fexingo