# Richard Sutton – Father of RL thinks LLMs are a dead end

Page: https://stenobird.com/podcast/dwarkesh-podcast/richard-sutton-father-of-rl-thinks-llms-are-a-dead-end
Text version: https://stenobird.com/podcast/dwarkesh-podcast/richard-sutton-father-of-rl-thinks-llms-are-a-dead-end.md
Podcast: [Dwarkesh Podcast](https://stenobird.com/podcast/dwarkesh-podcast)
Published: 2025-09-26T14:48:59+00:00
Episode link: https://www.dwarkesh.com/p/richard-sutton
Audio file: https://api.substack.com/feed/podcast/174609513/9ef910c6173bfc95dcec01b4b5ff7201.mp3
Processing state: not_requested
JSON: https://stenobird.com/v1/public/podcasts/dwarkesh-podcast/episodes/richard-sutton-father-of-rl-thinks-llms-are-a-dead-end
Duration seconds: 3982

## Resource

Richard Sutton is the father of reinforcement learning, winner of the 2024 Turing Award, and author of The Bitter Lesson. And he thinks LLMs are a dead end. After interviewing him, my steel man of Richard’s position is this: LLMs aren’t capable of learning on-the-job, so no matter how much we scale, we’ll need some new architecture to enable continual learning. And once we have it, we won’t need a special training phase — the agent will just learn on-the-fly, like all humans, and indeed, like all animals. This new paradigm will render our current approach with LLMs obsolete. In our interview, I did my best to represent the view that LLMs might function as the foundation on which experiential learning can happen… Some sparks flew. A big thanks to the Alberta Machine Intelligence Institute for inviting me up to Edmonton and for letting me use their studio and equipment. Enjoy! Watch on YouTube ; listen on Apple Podcasts or Spotify . Sponsors * Labelbox makes it possible to train AI agents in hyperrealistic RL environments. With an experienced team of applied researchers and a massive network of subject-matter experts, Labelbox ensures your training reflects important, real-world nuance. Turn your demo projects into working systems at labelbox.com/dwarkesh * Gemini Deep Research is designed for thorough exploration of hard topics. For this episode, it helped me trace reinforcement learning from early policy gradients up to current-day methods, combining clear explanations with curated examples. Try it out yourself at gemini.google.com * Hudson River Trading doesn’t silo their teams. Instead, HRT researchers openly trade ideas and share strategy code in a mono-repo. This means you’re able to learn at incredible speed and your contributions have impact across the entire fir…

## Actions

- request_transcript: `POST https://stenobird.com/v1/public/podcasts/dwarkesh-podcast/episodes/richard-sutton-father-of-rl-thinks-llms-are-a-dead-end/transcription-requests` — Idempotently request low-priority transcript generation for this episode.
- read_markdown: `GET https://stenobird.com/podcast/dwarkesh-podcast/richard-sutton-father-of-rl-thinks-llms-are-a-dead-end.md` — Read the agent-friendly Markdown representation of this episode resource.

A page view does not enqueue transcription. Agents should invoke `request_transcript` explicitly when they need this episode processed.

## Transcript

Full transcripts are not published on public pages unless there is a clear rights basis.