Episode

"Astra is much better at reasoning with filler tokens than previous models" by Dylan Xu, SebastianP, Alek Westover

Podcast
LessWrong (Curated & Popular)
Published
Sep 12, 2026
Duration seconds
933
Processing state
not_requested
Canonical source
https://www.buzzsprout.com/2037297/episodes/19793932-astra-is-much-better-at-reasoning-with-filler-tokens-than-previous-models-by-dylan-xu-sebastianp-alek-westover.mp3
Audio
https://www.buzzsprout.com/2037297/episodes/19793932-astra-is-much-better-at-reasoning-with-filler-tokens-than-previous-models-by-dylan-xu-sebastianp-alek-westover.mp3
JSON
/v1/public/podcasts/lesswrong-curated-popular-5643401/episodes/astra-is-much-better-at-reasoning-with-filler-tokens-than-previous-models-by-dylan-xu-sebastianp-alek-westover
Markdown
/podcast/lesswrong-curated-popular-5643401/astra-is-much-better-at-reasoning-with-filler-tokens-than-previous-models-by-dylan-xu-sebastianp-alek-westover.md

Actions

  • POST https://stenobird.com/v1/public/podcasts/lesswrong-curated-popular-5643401/episodes/astra-is-much-better-at-reasoning-with-filler-tokens-than-previous-models-by-dylan-xu-sebastianp-alek-westover/transcription-requests
    Idempotently request low-priority transcript generation for this episode.
  • GET https://stenobird.com/podcast/lesswrong-curated-popular-5643401/astra-is-much-better-at-reasoning-with-filler-tokens-than-previous-models-by-dylan-xu-sebastianp-alek-westover.md
    Read the agent-friendly Markdown representation of this episode resource.

Summary

We measure GPT-6-Astra's capabilities when its prompt is padded with a variable number of meaningless “filler” tokens (e.g., dots) and it is told to answer immediately without reasoning. On tasks designed to require lots of serial cognition, Astra performs significantly better with filler tokens than without (e.g., improving from ~10% to ~50% on 4-hop natural facts reasoning). On more general benchmarks, filler tokens also modestly improve Astra's performance (e.g., improving from ~60% to ~9...