# Confessions of a Large Language Model Page: https://stenobird.com/podcast/paul-weiss-waking-up-with-ai-6823804/confessions-of-a-large-language-model Text version: https://stenobird.com/podcast/paul-weiss-waking-up-with-ai-6823804/confessions-of-a-large-language-model.md Podcast: [Paul, Weiss Waking Up With AI](https://stenobird.com/podcast/paul-weiss-waking-up-with-ai-6823804) Published: 2026-01-22T19:31:51+00:00 Episode link: https://www.paulweiss.com/insights/podcasts/paul-weiss-waking-up-with-ai/ep-98-confessions-of-a-large-language-model Audio file: https://mcdn.podbean.com/mf/web/ubb7qn9df5p2ntxe/Ep_98_-_ConfessionsOfALargeLanguageModel8kl6o.mp3 Processing state: not_requested JSON: https://stenobird.com/v1/public/podcasts/paul-weiss-waking-up-with-ai-6823804/episodes/confessions-of-a-large-language-model Duration seconds: 1361 ## Resource In this episode, Katherine Forrest and Scott Caravello unpack OpenAI researchers’ proposed “confessions” framework designed to monitor for and detect dishonest outputs. They break down the researchers’ proof of concept results and the framework’s resilience to reward hacking, along with its limits in connection with hallucinations. Then they turn to Google DeepMind’s “Distributional AGI Safety,” exploring a hypothetical path to AGI via a patchwork of agents and routing infrastructure, as well as the authors’ proposed four layer safety stack. ## Learn More About Paul, Weiss’s Artificial Intelligence practice:https://www.paulweiss.com/industries/artificial-intelligence ## Actions - request_transcript: `POST https://stenobird.com/v1/public/podcasts/paul-weiss-waking-up-with-ai-6823804/episodes/confessions-of-a-large-language-model/transcription-requests` — Idempotently request low-priority transcript generation for this episode. - read_markdown: `GET https://stenobird.com/podcast/paul-weiss-waking-up-with-ai-6823804/confessions-of-a-large-language-model.md` — Read the agent-friendly Markdown representation of this episode resource. A page view does not enqueue transcription. Agents should invoke `request_transcript` explicitly when they need this episode processed. ## Transcript Full transcripts are not published on public pages unless there is a clear rights basis.