# #367 Don't Build on Jell-O: How to Make Agentic AI Reliable with Dan Klein, CTO at Scaled Cognition Page: https://stenobird.com/podcast/dataframed/367-don-t-build-on-jell-o-how-to-make-agentic-ai-reliable-with-dan-klein-cto-at-scaled-cognition Text version: https://stenobird.com/podcast/dataframed/367-don-t-build-on-jell-o-how-to-make-agentic-ai-reliable-with-dan-klein-cto-at-scaled-cognition.md Podcast: [DataFramed](https://stenobird.com/podcast/dataframed) Published: 2026-07-06T09:00:00+00:00 Episode link: https://www.datacamp.com/podcast Audio file: https://dts.podtrac.com/redirect.mp3/cohst.app/pdcst/6G1A6D/episodes.captivate.fm/episode/ae3fe234-a750-4de3-9ada-1e980c3c7744.mp3 Processing state: not_requested JSON: https://stenobird.com/v1/public/podcasts/dataframed/episodes/367-don-t-build-on-jell-o-how-to-make-agentic-ai-reliable-with-dan-klein-cto-at-scaled-cognition Duration seconds: 3102 ## Resource Across the AI industry, capability has exploded while trustworthiness has lagged badly behind. The same technology that writes fluent prose can invent a refund policy that was never real, and most of those errors are subtle enough that no one notices. As more teams hand high-stakes work to AI — in banking, healthcare, customer service — the cost of confident mistakes adds up fast. So how common are hallucinations, really? Can chaining models together or adding humans to the loop fix it? And is reliability something you can design into a system from the start? Dan Klein is the CTO and co-founder of Scaled Cognition and a professor of computer science at UC Berkeley, where he leads the Berkeley NLP Group within the Berkeley AI Research (BAIR) Lab. He previously co-founded Semantic Machines, a conversational AI company acquired by Microsoft in 2018. At Scaled Cognition he built APT (Agentic Pretrained Transformer), a frontier model designed from the ground up for reliable, policy-adherent agentic AI. In the episode, Richie and Dan explore why AI reliability has lagged behind capability, how hallucinations hide in plain sight, the limits of humans-in-the-loop and LLM-as-judge, building reliability into model architecture, agentic systems and verifiable actions, test-driven agent development, the skills that stay valuable, digital literacy, and much more. Links Mentioned in the Show: • Connect with Dan: https://www.linkedin.com/in/dan-klein/ • Scaled Cognition: https://www.scaledcognition.com/ • Berkeley NLP Group: https://nlp.cs.berkeley.edu/ • Code smells (Martin Fowler): https://martinfowler.com/bliki/CodeSmell.html • Refactoring, by Martin Fowler: https://martinfowler.com/books/refactoring.html • "Now you have two problems" (Jamie Zawinski quote): https://regex.info/blo… ## Actions - request_transcript: `POST https://stenobird.com/v1/public/podcasts/dataframed/episodes/367-don-t-build-on-jell-o-how-to-make-agentic-ai-reliable-with-dan-klein-cto-at-scaled-cognition/transcription-requests` — Idempotently request low-priority transcript generation for this episode. - read_markdown: `GET https://stenobird.com/podcast/dataframed/367-don-t-build-on-jell-o-how-to-make-agentic-ai-reliable-with-dan-klein-cto-at-scaled-cognition.md` — Read the agent-friendly Markdown representation of this episode resource. A page view does not enqueue transcription. Agents should invoke `request_transcript` explicitly when they need this episode processed. ## Transcript Full transcripts are not published on public pages unless there is a clear rights basis.