Episode
"Can activation verbalizers surface an internal chain of thought?" by oakhu, ryan_greenblatt
- Published
- Jun 22, 2026
- Duration seconds
- 4778
- Processing state
not_requested
Actions
POST https://stenobird.com/v1/public/podcasts/lesswrong-curated-popular-5643401/episodes/can-activation-verbalizers-surface-an-internal-chain-of-thought-by-oakhu-ryan-greenblatt/transcription-requests
Idempotently request low-priority transcript generation for this episode.GET https://stenobird.com/podcast/lesswrong-curated-popular-5643401/can-activation-verbalizers-surface-an-internal-chain-of-thought-by-oakhu-ryan-greenblatt.md
Read the agent-friendly Markdown representation of this episode resource.
Summary
We introduce an evaluation for activation verbalizers: can they surface a target model's reasoning as it solves a math problem in a single forward pass? For open-weight NLAs, the answer seems to be: "possibly, but definitely not reliably". Lots of important capabilities currently require AI models to reason "out loud" in a natural-language chain of thought, which means that we can monitor important parts of their thinking. It would be nice to have this same affordance for the reasoning tha...