{"podcast":{"title":"Ship It Weekly - DevOps, SRE, Platform and Cloud Engineering News","slug":"ship-it-weekly-devops-sre-platform-and-cloud-engineering-news-7591275","podcast_index_feed_id":7591275,"rss_url":"https://media.rss.com/ship-it-weekly/feed.xml","website_url":"https://www.shipitweekly.fm/","image_url":"https://media.rss.com/ship-it-weekly/20260103_010109_fc16278a46c7b2c61123ed668a34f79d.jpg","author":"Teller's Tech - DevOps, SRE and Cloud Podcast","episode_count":55,"summary":"Ship It Weekly is a short, practical recap of what actually matters in DevOps, SRE, cloud infrastructure, and platform engineering. Each episode, your host Brian Teller walks through the latest outages, releases, tools, and incident writeups, then translates them into “here’s what this means for your systems” instead of just reading headlines. Expect a couple of main stories with context, a quick hit of tools or releases worth bookmarking, and the occasional segment on on-call, burnout, or team culture. This isn’t a certification prep show or a lab walkthrough. It’s aimed at people who are already working in the space and want to stay sharp without scrolling status pages, cloud updates, and blogs all week. You’ll hear about things like cloud provider incidents, Kubernetes and platform trends, Terraform and infrastructure changes, and real postmortems that are actually worth your time. Most episodes are 15–30 minutes, so you can catch up on the way to work or between meetings. Every now and then there will be a “special” focused on a big outage or a specific theme, but the default format is simple: what happened, why it matters, and what you might want to do about it in your own en…","last_synced_at":"2026-07-19T00:20:39.423055+00:00","page_url":"https://stenobird.com/podcast/ship-it-weekly-devops-sre-platform-and-cloud-engineering-news-7591275"},"episode":{"title":"Ship It Conversations: Meta’s Francois Richard on AI Incident Response, SLOs, and Reliability at Scale","slug":"ship-it-conversations-meta-s-francois-richard-on-ai-incident-response-slos-and-reliability-at-scale","published_at":"2026-06-16T02:06:13+00:00","page_url":"https://stenobird.com/podcast/ship-it-weekly-devops-sre-platform-and-cloud-engineering-news-7591275/ship-it-conversations-meta-s-francois-richard-on-ai-incident-response-slos-and-reliability-at-scale","show_page_url":"https://stenobird.com/podcast/ship-it-weekly-devops-sre-platform-and-cloud-engineering-news-7591275","url":"https://rss.com/podcasts/ship-it-weekly/2918430","audio_url":"https://content.rss.com/episodes/356364/2918430/ship-it-weekly/2026_06_16_01_51_37_510a243d-0dca-47ee-976e-afb3cd23636c.mp3","summary":"This is a guest conversation episode of Ship It Weekly , separate from the weekly news recaps. In this Ship It: Conversations episode, I talk with Francois Richard, Engineering Director at Meta, about reliability at scale, how AI is changing production risk, what teams actually learn from incidents, and why recovery practice matters just as much as prevention. We talk about the proactive and reactive sides of reliability, why SLOs should represent a promise to users instead of just another dashboard number, how incident reviews should drive real system improvements, and how teams can practice recovery before production forces the lesson on them. The bigger theme here is that reliability is not just about avoiding failure. It is about knowing what happens when prevention fails. That means practicing regional failure, understanding overload behavior, improving incident response, using AI carefully during investigation, and making reliability targets match the actual lifecycle and importance of the system. Highlights • Why reliability work starts with both prevention and recovery • The difference between reactive incident response and proactive reliability engineering • How Meta thinks about disaster recovery testing and regional failure practice • Why an SLO should be treated like a promise to users, not just a dashboard metric • How SLO trends help teams decide when to invest more in reliability or take more product risk • What engineers actually learn during the “pressure cooker” of an incident • Why incident reviews should produce follow-up work, not just a nicer explanation of what broke • The difference between finding the cause of an incident and improving the system • Where AI agents can help with incident investigation, telemetry, metrics, and query building • Wh…","meta_description":"This is a guest conversation episode of Ship It Weekly , separate from the weekly news recaps. In this Ship It: Conversations episode, I talk with Francoi…","key_points":[],"chapters":[],"topics":[],"duration_seconds":2576,"processing_state":"not_requested","actions":[{"name":"request_transcript","method":"POST","url":"https://stenobird.com/v1/public/podcasts/ship-it-weekly-devops-sre-platform-and-cloud-engineering-news-7591275/episodes/ship-it-conversations-meta-s-francois-richard-on-ai-incident-response-slos-and-reliability-at-scale/transcription-requests","description":"Idempotently request low-priority transcript generation for this episode."},{"name":"read_markdown","method":"GET","url":"https://stenobird.com/podcast/ship-it-weekly-devops-sre-platform-and-cloud-engineering-news-7591275/ship-it-conversations-meta-s-francois-richard-on-ai-incident-response-slos-and-reliability-at-scale.md","description":"Read the agent-friendly Markdown representation of this episode resource."}]}}