{"podcast":{"title":"The Site Reliability Podcast with Fexingo: SRE, Uptime, and Production Engineering","slug":"the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923","podcast_index_feed_id":7871923,"rss_url":"https://feeds.fexingo.com/business/the-site-reliability-podcast.xml","website_url":"https://www.fexingo.com/","image_url":"https://audio.fexingo.com/business/the-site-reliability-podcast/cover.png","author":"Fexingo","episode_count":119,"summary":"Lucas and Luna cut through the noise around site reliability engineering to examine how real-world SRE teams balance uptime, incident response, and production change. Each episode takes a single concept — error budgets, toil automation, postmortem culture, capacity planning — and grounds it in a specific case: how a major streaming service reduced paging noise, how a payments platform rebuilt its incident command structure, or how a cloud provider manages multi-region failover. Lucas brings the numbers — latency percentiles, MTTR trends, SLO burn rates — while Luna pushes on the human and organizational trade-offs: What does a junior SRE need to know about on-call? How do you measure reliability without crushing innovation? Why do some blameless postmortems actually work? Together they treat SRE not as a certification topic but as a living practice, citing real outages, open-source tools, and engineering blogs. This show is for engineers, ops leads, and platform teams who already know the basics and want to debate the hard edges: Is 99.999% uptime always worth the cost? When should you deliberately degrade service to improve reliability? How do you design for resilience when your…","last_synced_at":"2026-07-18T22:18:41.635042+00:00","page_url":"https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923"},"episode":{"title":"How SRE Teams Use AI for Incident Triage and Root Cause Analysis","slug":"how-sre-teams-use-ai-for-incident-triage-and-root-cause-analysis","published_at":"2026-06-24T22:21:14+00:00","page_url":"https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-ai-for-incident-triage-and-root-cause-analysis","show_page_url":"https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923","url":"https://audio.fexingo.com/business/the-site-reliability-podcast/episode-0071.mp3","audio_url":"https://audio.fexingo.com/business/the-site-reliability-podcast/episode-0071.mp3","summary":"Episode 71 of The Site Reliability Podcast with Fexingo dives into how SRE teams are applying large language models and AI assistants to accelerate incident triage and root cause analysis. Lucas and Luna examine a real case from a mid-sized e-commerce platform: after a database connection pool exhaustion caused a 14-minute partial outage, the on-call engineer used a locally-run AI tool to correlate error logs, recent deployments, and metric changes in under 90 seconds — cutting mean time to resolution by 40% compared to manual investigation. The hosts break down the technical setup (RAG pipeline fed by runbooks and past postmortems), the risks (hallucination, over-reliance, training data staleness), and the operational guardrails (human-in-the-loop verification, confidence scoring, and mandatory peer review for any AI-generated remediation). They also discuss why some teams are starting with 'read-only' AI triage before granting write access. No hype, just the engineering reality of LLMs in the incident response workflow as of mid-2026. #SiteReliabilityEngineering #IncidentResponse #AIOps #LLM #RootCauseAnalysis #IncidentTriage #RAG #OnCall #Postmortem #SRE #ProductionEngineering #Uptime #Technology #FexingoBusiness #BusinessPodcast #AIAssistant #MeanTimeToResolution #HumanInTheLoop Keep every episode free: buymeacoffee.com/fexingo","meta_description":"Episode 71 of The Site Reliability Podcast with Fexingo dives into how SRE teams are applying large language models and AI assistants to accelerate incide…","key_points":[],"chapters":[],"topics":[],"duration_seconds":662,"processing_state":"not_requested","actions":[{"name":"request_transcript","method":"POST","url":"https://stenobird.com/v1/public/podcasts/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/episodes/how-sre-teams-use-ai-for-incident-triage-and-root-cause-analysis/transcription-requests","description":"Idempotently request low-priority transcript generation for this episode."},{"name":"read_markdown","method":"GET","url":"https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-ai-for-incident-triage-and-root-cause-analysis.md","description":"Read the agent-friendly Markdown representation of this episode resource."}]}}