{"podcast":{"title":"Dwarkesh Podcast","slug":"dwarkesh-podcast","podcast_index_feed_id":1251189,"rss_url":"https://api.substack.com/feed/podcast/69345.rss","website_url":"https://www.dwarkesh.com/podcast","image_url":"https://substackcdn.com/feed/podcast/69345/db6ef1755a45c6e0e7a478f6dbe68984.jpg","author":"Dwarkesh Patel","episode_count":128,"summary":"Deeply researched interviews","last_synced_at":"2026-06-05T06:19:20.066984+00:00","page_url":"https://stenobird.com/podcast/dwarkesh-podcast"},"episode":{"title":"Richard Sutton – Father of RL thinks LLMs are a dead end","slug":"richard-sutton-father-of-rl-thinks-llms-are-a-dead-end","published_at":"2025-09-26T14:48:59+00:00","page_url":"https://stenobird.com/podcast/dwarkesh-podcast/richard-sutton-father-of-rl-thinks-llms-are-a-dead-end","show_page_url":"https://stenobird.com/podcast/dwarkesh-podcast","url":"https://www.dwarkesh.com/p/richard-sutton","audio_url":"https://api.substack.com/feed/podcast/174609513/9ef910c6173bfc95dcec01b4b5ff7201.mp3","summary":"Richard Sutton is the father of reinforcement learning, winner of the 2024 Turing Award, and author of The Bitter Lesson. And he thinks LLMs are a dead end. After interviewing him, my steel man of Richard’s position is this: LLMs aren’t capable of learning on-the-job, so no matter how much we scale, we’ll need some new architecture to enable continual learning. And once we have it, we won’t need a special training phase — the agent will just learn on-the-fly, like all humans, and indeed, like all animals. This new paradigm will render our current approach with LLMs obsolete. In our interview, I did my best to represent the view that LLMs might function as the foundation on which experiential learning can happen… Some sparks flew. A big thanks to the Alberta Machine Intelligence Institute for inviting me up to Edmonton and for letting me use their studio and equipment. Enjoy! Watch on YouTube ; listen on Apple Podcasts or Spotify . Sponsors * Labelbox makes it possible to train AI agents in hyperrealistic RL environments. With an experienced team of applied researchers and a massive network of subject-matter experts, Labelbox ensures your training reflects important, real-world nuance. Turn your demo projects into working systems at labelbox.com/dwarkesh * Gemini Deep Research is designed for thorough exploration of hard topics. For this episode, it helped me trace reinforcement learning from early policy gradients up to current-day methods, combining clear explanations with curated examples. Try it out yourself at gemini.google.com * Hudson River Trading doesn’t silo their teams. Instead, HRT researchers openly trade ideas and share strategy code in a mono-repo. This means you’re able to learn at incredible speed and your contributions have impact across the entire fir…","meta_description":"Richard Sutton is the father of reinforcement learning, winner of the 2024 Turing Award, and author of The Bitter Lesson. And he thinks LLMs are a dead en…","key_points":[],"chapters":[],"topics":[],"duration_seconds":3982,"processing_state":"not_requested","actions":[{"name":"request_transcript","method":"POST","url":"https://stenobird.com/v1/public/podcasts/dwarkesh-podcast/episodes/richard-sutton-father-of-rl-thinks-llms-are-a-dead-end/transcription-requests","description":"Idempotently request low-priority transcript generation for this episode."},{"name":"read_markdown","method":"GET","url":"https://stenobird.com/podcast/dwarkesh-podcast/richard-sutton-father-of-rl-thinks-llms-are-a-dead-end.md","description":"Read the agent-friendly Markdown representation of this episode resource."}]}}