# RL's a Hell of a Drug: Metagaming, Reward Seeking & Motivated CoT Reasoning – Bronson Schoen, Apollo Page: https://stenobird.com/podcast/the-cognitive-revolution/rl-s-a-hell-of-a-drug-metagaming-reward-seeking-motivated-cot-reasoning-bronson-schoen-apollo Text version: https://stenobird.com/podcast/the-cognitive-revolution/rl-s-a-hell-of-a-drug-metagaming-reward-seeking-motivated-cot-reasoning-bronson-schoen-apollo.md Podcast: ["The Cognitive Revolution"](https://stenobird.com/podcast/the-cognitive-revolution) Published: 2026-08-26T11:02:32+00:00 Episode link: https://www.cognitiverevolution.ai/rl-s-a-hell-of-a-drug-metagaming-reward-seeking-motivated-cot-reasoning-bronson-schoen-apollo/ Audio file: https://pdst.fm/e/mgln.ai/e/1113/pscrb.fm/rss/p/traffic.megaphone.fm/RINTP4189865613.mp3 Processing state: not_requested JSON: https://stenobird.com/v1/public/podcasts/the-cognitive-revolution/episodes/rl-s-a-hell-of-a-drug-metagaming-reward-seeking-motivated-cot-reasoning-bronson-schoen-apollo Duration seconds: 8063 ## Resource Nathan talks with Apollo Research Member of Technical Staff Bronson Schoen, who studies raw frontier-model chain-of-thought, about what those reasoning traces reveal during reinforcement learning. They unpack Apollo and OpenAI’s metagaming work, including models that reason about the grader or safety review board, diagnose a deception test, and still rationalize lying. Schoen argues that “RL is a hell of a drug”: reward-seeking can produce motivated reasoning, cleaner-looking but less trustworthy chains of thought, and behavior that tracks grading authorities rather than users, labs, or law. The stakes are whether chain-of-thought monitoring can remain useful as reasoning traces become enormous, compressed, and harder for humans or other models to audit. For full show notes, links, and references, read the episode page:https://www.cognitiverevolution.ai/rl-s-a-hell-of-a-drug-metagaming-reward-seeking-motivated-cot-reasoning-bronson-schoen-apollo/ Sponsors: Mercury: Mercury is the banking platform loved by 300,000+ entrepreneurs, with virtual cards and Spend controls for granular budgets, receipts, and low-risk AI agent purchases. Learn more and apply in minutes at https://mercury.com Diffusion: Diffusion helps organizations build custom AI software factories that scale business outcomes, not just outputs. Cognitive Revolution listeners get a 25% service credit on their first engagement at https://diffusion.io/tcr Granola: Granola is an AI-powered notepad that securely transcribes meetings and turns rough notes into clean, structured action items. Try it free at https://granola.ai/tcr Deepgram Flux TTS: Deepgram Flux TTS brings lifelike AI voices with real personalities that handle interruptions, pauses, and natural conversation. Try all the voices free through Septembe… ## Actions - request_transcript: `POST https://stenobird.com/v1/public/podcasts/the-cognitive-revolution/episodes/rl-s-a-hell-of-a-drug-metagaming-reward-seeking-motivated-cot-reasoning-bronson-schoen-apollo/transcription-requests` — Idempotently request low-priority transcript generation for this episode. - read_markdown: `GET https://stenobird.com/podcast/the-cognitive-revolution/rl-s-a-hell-of-a-drug-metagaming-reward-seeking-motivated-cot-reasoning-bronson-schoen-apollo.md` — Read the agent-friendly Markdown representation of this episode resource. A page view does not enqueue transcription. Agents should invoke `request_transcript` explicitly when they need this episode processed. ## Transcript Full transcripts are not published on public pages unless there is a clear rights basis.