{"podcast":{"title":"\"The Cognitive Revolution\"","slug":"the-cognitive-revolution","podcast_index_feed_id":6011783,"rss_url":"https://feeds.megaphone.fm/RINTP3108857801","website_url":"https://www.cognitiverevolution.ai/","image_url":"https://megaphone.imgix.net/podcasts/30f818da-c930-11ed-9b4b-1352ca96fb17/image/888e2c534b7c2534213c97e025646932.png?ixlib=rails-4.3.1&max-w=3000&max-h=3000&fit=crop&auto=format,compress","author":"Turpentine","episode_count":376,"summary":"A biweekly podcast where hosts Nathan Labenz and Erik Torenberg interview the builders on the edge of AI and explore the dramatic shift it will unlock in the coming years. The Cognitive Revolution is part of the Turpentine podcast network. To learn more: turpentine.co","last_synced_at":"2026-09-20T18:17:39.232965+00:00","page_url":"https://stenobird.com/podcast/the-cognitive-revolution"},"episode":{"title":"RL's a Hell of a Drug: Metagaming, Reward Seeking & Motivated CoT Reasoning – Bronson Schoen, Apollo","slug":"rl-s-a-hell-of-a-drug-metagaming-reward-seeking-motivated-cot-reasoning-bronson-schoen-apollo","published_at":"2026-08-26T11:02:32+00:00","page_url":"https://stenobird.com/podcast/the-cognitive-revolution/rl-s-a-hell-of-a-drug-metagaming-reward-seeking-motivated-cot-reasoning-bronson-schoen-apollo","show_page_url":"https://stenobird.com/podcast/the-cognitive-revolution","url":"https://www.cognitiverevolution.ai/rl-s-a-hell-of-a-drug-metagaming-reward-seeking-motivated-cot-reasoning-bronson-schoen-apollo/","audio_url":"https://pdst.fm/e/mgln.ai/e/1113/pscrb.fm/rss/p/traffic.megaphone.fm/RINTP4189865613.mp3","summary":"Nathan talks with Apollo Research Member of Technical Staff Bronson Schoen, who studies raw frontier-model chain-of-thought, about what those reasoning traces reveal during reinforcement learning. They unpack Apollo and OpenAI’s metagaming work, including models that reason about the grader or safety review board, diagnose a deception test, and still rationalize lying. Schoen argues that “RL is a hell of a drug”: reward-seeking can produce motivated reasoning, cleaner-looking but less trustworthy chains of thought, and behavior that tracks grading authorities rather than users, labs, or law. The stakes are whether chain-of-thought monitoring can remain useful as reasoning traces become enormous, compressed, and harder for humans or other models to audit. For full show notes, links, and references, read the episode page:https://www.cognitiverevolution.ai/rl-s-a-hell-of-a-drug-metagaming-reward-seeking-motivated-cot-reasoning-bronson-schoen-apollo/ Sponsors: Mercury: Mercury is the banking platform loved by 300,000+ entrepreneurs, with virtual cards and Spend controls for granular budgets, receipts, and low-risk AI agent purchases. Learn more and apply in minutes at https://mercury.com Diffusion: Diffusion helps organizations build custom AI software factories that scale business outcomes, not just outputs. Cognitive Revolution listeners get a 25% service credit on their first engagement at https://diffusion.io/tcr Granola: Granola is an AI-powered notepad that securely transcribes meetings and turns rough notes into clean, structured action items. Try it free at https://granola.ai/tcr Deepgram Flux TTS: Deepgram Flux TTS brings lifelike AI voices with real personalities that handle interruptions, pauses, and natural conversation. Try all the voices free through Septembe…","meta_description":"Nathan talks with Apollo Research Member of Technical Staff Bronson Schoen, who studies raw frontier-model chain-of-thought, about what those reasoning tr…","key_points":[],"chapters":[],"topics":[],"duration_seconds":8063,"processing_state":"not_requested","actions":[{"name":"request_transcript","method":"POST","url":"https://stenobird.com/v1/public/podcasts/the-cognitive-revolution/episodes/rl-s-a-hell-of-a-drug-metagaming-reward-seeking-motivated-cot-reasoning-bronson-schoen-apollo/transcription-requests","description":"Idempotently request low-priority transcript generation for this episode."},{"name":"read_markdown","method":"GET","url":"https://stenobird.com/podcast/the-cognitive-revolution/rl-s-a-hell-of-a-drug-metagaming-reward-seeking-motivated-cot-reasoning-bronson-schoen-apollo.md","description":"Read the agent-friendly Markdown representation of this episode resource."}]}}