Episode

"Can activation verbalizers surface an internal chain of thought?" by oakhu, ryan_greenblatt

Podcast
LessWrong (Curated & Popular)
Published
Jun 22, 2026
Duration seconds
4778
Processing state
not_requested
Canonical source
https://www.buzzsprout.com/2037297/episodes/19382397-can-activation-verbalizers-surface-an-internal-chain-of-thought-by-oakhu-ryan_greenblatt.mp3
Audio
https://www.buzzsprout.com/2037297/episodes/19382397-can-activation-verbalizers-surface-an-internal-chain-of-thought-by-oakhu-ryan_greenblatt.mp3
JSON
/v1/public/podcasts/lesswrong-curated-popular-5643401/episodes/can-activation-verbalizers-surface-an-internal-chain-of-thought-by-oakhu-ryan-greenblatt
Markdown
/podcast/lesswrong-curated-popular-5643401/can-activation-verbalizers-surface-an-internal-chain-of-thought-by-oakhu-ryan-greenblatt.md

Actions

  • POST https://stenobird.com/v1/public/podcasts/lesswrong-curated-popular-5643401/episodes/can-activation-verbalizers-surface-an-internal-chain-of-thought-by-oakhu-ryan-greenblatt/transcription-requests
    Idempotently request low-priority transcript generation for this episode.
  • GET https://stenobird.com/podcast/lesswrong-curated-popular-5643401/can-activation-verbalizers-surface-an-internal-chain-of-thought-by-oakhu-ryan-greenblatt.md
    Read the agent-friendly Markdown representation of this episode resource.

Summary

We introduce an evaluation for activation verbalizers: can they surface a target model's reasoning as it solves a math problem in a single forward pass? For open-weight NLAs, the answer seems to be: "possibly, but definitely not reliably". Lots of important capabilities currently require AI models to reason "out loud" in a natural-language chain of thought, which means that we can monitor important parts of their thinking. It would be nice to have this same affordance for the reasoning tha...