# The Site Reliability Podcast with Fexingo: SRE, Uptime, and Production Engineering Page: https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923 Text version: https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923.md RSS feed: https://feeds.fexingo.com/business/the-site-reliability-podcast.xml Official site: https://www.fexingo.com/ Author: Fexingo Episodes: 119 ## Resource Lucas and Luna cut through the noise around site reliability engineering to examine how real-world SRE teams balance uptime, incident response, and production change. Each episode takes a single concept — error budgets, toil automation, postmortem culture, capacity planning — and grounds it in a specific case: how a major streaming service reduced paging noise, how a payments platform rebuilt its incident command structure, or how a cloud provider manages multi-region failover. Lucas brings the numbers — latency percentiles, MTTR trends, SLO burn rates — while Luna pushes on the human and organizational trade-offs: What does a junior SRE need to know about on-call? How do you measure reliability without crushing innovation? Why do some blameless postmortems actually work? Together they treat SRE not as a certification topic but as a living practice, citing real outages, open-source tools, and engineering blogs. This show is for engineers, ops leads, and platform teams who already know the basics and want to debate the hard edges: Is 99.999% uptime always worth the cost? When should you deliberately degrade service to improve reliability? How do you design for resilience when your… ## Machine-readable JSON: https://stenobird.com/v1/public/podcasts/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923 Markdown: https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923.md ## Episodes - [How SRE Teams Use Observability Signals to Diagnose Production Issues](https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-observability-signals-to-diagnose-production-issues) — 2026-07-18T22:07:04+00:00 - [How SRE Teams Use Burn Rate Alerts to Stay Within Error Budgets](https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-burn-rate-alerts-to-stay-within-error-budgets) — 2026-07-18T09:52:08+00:00 - [How Slack Cut Mean Time to Acknowledge by 60 Percent With On-Call Orchestration](https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-slack-cut-mean-time-to-acknowledge-by-60-percent-with-on-call-orchestration) — 2026-07-17T22:37:32+00:00 - [How SRE Teams Use Service Level Objectives to Align Business and Engineering](https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-service-level-objectives-to-align-business-and-engineering) — 2026-07-17T10:12:25+00:00 - [How SRE Teams Use Load Shedding to Protect Critical Services](https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-load-shedding-to-protect-critical-services) — 2026-07-16T22:32:03+00:00 - [How SRE Teams Use Toil Budgets to Automate the Right Things](https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-toil-budgets-to-automate-the-right-things) — 2026-07-16T10:12:41+00:00 - [How SRE Teams Use Fault Trees to Root Out Latent Defects](https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-fault-trees-to-root-out-latent-defects) — 2026-07-15T22:19:40+00:00 - [How SRE Teams Use Game Days to Build Incident Muscle Memory](https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-game-days-to-build-incident-muscle-memory) — 2026-07-15T10:08:28+00:00 - [How SRE Teams Use Error Budgets to Balance Reliability and Velocity](https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-error-budgets-to-balance-reliability-and-velocity-3) — 2026-07-14T22:16:31+00:00 - [How SRE Teams Use Incident Metrics to Improve Postmortem Quality](https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-incident-metrics-to-improve-postmortem-quality) — 2026-07-14T10:06:23+00:00 - [How SRE Teams Use Chaos Engineering to Find Hidden Failure Modes](https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-chaos-engineering-to-find-hidden-failure-modes) — 2026-07-13T22:47:11+00:00 - [How SRE Teams Use Latency SLOs to Improve User Experience](https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-latency-slos-to-improve-user-experience) — 2026-07-13T10:37:53+00:00 - [How SRE Teams Use Incident Postmortems for Systemic Improvement](https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-incident-postmortems-for-systemic-improvement) — 2026-07-12T22:43:38+00:00 - [How SRE Teams Use Capacity Planning to Prevent Outages Before They Happen](https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-capacity-planning-to-prevent-outages-before-they-happen) — 2026-07-12T10:08:46+00:00 - [How SRE Teams Use Non-Abstract Large System Design to Prevent Outages](https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-non-abstract-large-system-design-to-prevent-outages) — 2026-07-11T22:28:13+00:00 - [How SRE Teams Use Runbooks to Standardize Incident Response](https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-runbooks-to-standardize-incident-response) — 2026-07-11T10:05:17+00:00 - [How Airbnb Uses Traffic Shifting to Prevent Cascading Failures](https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-airbnb-uses-traffic-shifting-to-prevent-cascading-failures) — 2026-07-10T22:37:50+00:00 - [How SRE Teams Manage Cognitive Load During Incidents](https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-manage-cognitive-load-during-incidents) — 2026-07-10T10:34:39+00:00 - [How SRE Teams Use Dependency Graphs to Prevent Cascading Failures](https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-dependency-graphs-to-prevent-cascading-failures) — 2026-07-09T22:26:50+00:00 - [How SRE Teams Use Canary Deployments to Reduce Release Risk](https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-canary-deployments-to-reduce-release-risk) — 2026-07-09T10:08:43+00:00 - [How SRE Teams Use Incident Cost Metrics to Justify Reliability Investment](https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-incident-cost-metrics-to-justify-reliability-investment) — 2026-07-08T22:27:45+00:00 - [How SRE Teams Use Observability Pipelines to Reduce Data Costs](https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-observability-pipelines-to-reduce-data-costs) — 2026-07-08T12:09:31+00:00 - [How SRE Teams Use Blameless Culture to Improve Incident Response](https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-blameless-culture-to-improve-incident-response-2) — 2026-07-07T22:28:24+00:00 - [How SRE Teams Use Incident Severity Frameworks to Triage Faster](https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-incident-severity-frameworks-to-triage-faster) — 2026-07-07T10:21:29+00:00 - [How SRE Teams Use Synthetic Monitoring to Catch Problems Before Users Do](https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-synthetic-monitoring-to-catch-problems-before-users-do) — 2026-07-06T22:25:09+00:00 - [How SRE Teams Use Error Budgets to Balance Reliability and Velocity](https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-error-budgets-to-balance-reliability-and-velocity) — 2026-07-04T10:27:23+00:00 - [How SRE Teams Use Incident Metrics to Improve Response](https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-incident-metrics-to-improve-response) — 2026-07-03T22:33:59+00:00 - [How SRE Teams Use Cost Optimization to Reduce Cloud Waste](https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-cost-optimization-to-reduce-cloud-waste) — 2026-07-03T10:22:07+00:00 - [How SRE Teams Use Toil Budgets to Protect Engineering Time](https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-toil-budgets-to-protect-engineering-time) — 2026-07-02T22:25:58+00:00 - [How SRE Teams Use Structured Fails to Learn Faster](https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-structured-fails-to-learn-faster) — 2026-07-02T10:26:55+00:00 - [How SRE Teams Use Post-Incident Reviews for System Improvements](https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-post-incident-reviews-for-system-improvements) — 2026-07-01T22:28:56+00:00 - [How SRE Teams Use Capacity Planning to Prevent Outages](https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-capacity-planning-to-prevent-outages) — 2026-07-01T12:47:58+00:00 - [How SRE Teams Use Chaos Engineering to Build Resilient Systems](https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-chaos-engineering-to-build-resilient-systems) — 2026-06-30T22:31:33+00:00 - [How SRE Teams Use Cost of Delay to Prioritize Reliability Work](https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-cost-of-delay-to-prioritize-reliability-work) — 2026-06-30T10:18:15+00:00 - [How SRE Teams Use Latency Budgets to Meet Performance SLOs](https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-latency-budgets-to-meet-performance-slos) — 2026-06-29T22:27:49+00:00 - [How SRE Teams Use Runbooks to Streamline Incident Response](https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-runbooks-to-streamline-incident-response) — 2026-06-29T10:28:00+00:00 - [How SRE Teams Use Observability to Reduce Mean Time to Detect](https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-observability-to-reduce-mean-time-to-detect) — 2026-06-28T22:21:49+00:00 - [How SRE Teams Use Service Level Agreements to Set Expectations](https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-service-level-agreements-to-set-expectations) — 2026-06-28T09:59:59+00:00 - [How SRE Teams Use Canary Deployments to Reduce Risk](https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-canary-deployments-to-reduce-risk) — 2026-06-27T22:28:28+00:00 - [How SRE Teams Use DORA Metrics to Measure DevOps Performance](https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-dora-metrics-to-measure-devops-performance) — 2026-06-27T10:05:17+00:00 - [How SRE Teams Use Service Level Objectives to Drive Reliability](https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-service-level-objectives-to-drive-reliability) — 2026-06-26T22:14:06+00:00 - [How SRE Teams Use Blameless Culture to Improve Incident Response](https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-blameless-culture-to-improve-incident-response) — 2026-06-26T10:10:33+00:00 - [How SRE Teams Use Blameless Postmortems to Build Trust](https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-blameless-postmortems-to-build-trust) — 2026-06-25T22:23:48+00:00 - [How SRE Teams Use Fault Tree Analysis to Prevent Root Causes](https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-fault-tree-analysis-to-prevent-root-causes) — 2026-06-25T09:58:05+00:00 - [How SRE Teams Use AI for Incident Triage and Root Cause Analysis](https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-ai-for-incident-triage-and-root-cause-analysis) — 2026-06-24T22:21:14+00:00 - [How SRE Teams Use Game Days to Test Incident Response](https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-game-days-to-test-incident-response) — 2026-06-24T10:11:19+00:00 - [How SRE Teams Use Error Budgets to Balance Reliability and Velocity](https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-error-budgets-to-balance-reliability-and-velocity-2) — 2026-06-23T22:16:39+00:00 - [How SRE Teams Use Infrastructure as Code to Prevent Configuration Drift](https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-infrastructure-as-code-to-prevent-configuration-drift) — 2026-06-23T10:16:36+00:00 - [How SRE Teams Use Incident Response Playbooks](https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-incident-response-playbooks) — 2026-06-22T22:08:27+00:00 - [How SRE Teams Use Readiness Checks to Prevent Bad Deployments](https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-readiness-checks-to-prevent-bad-deployments) — 2026-06-22T09:59:32+00:00 ## Actions Episode pages expose an explicit `request_transcript` action. A page view does not automatically enqueue transcription.