# Rethinking On-Policy Distillation of Large Language Models II: One Training Example Page: https://stenobird.com/podcast/daily-paper-cast-7079649/rethinking-on-policy-distillation-of-large-language-models-ii-one-training-example Text version: https://stenobird.com/podcast/daily-paper-cast-7079649/rethinking-on-policy-distillation-of-large-language-models-ii-one-training-example.md Podcast: [Daily Paper Cast](https://stenobird.com/podcast/daily-paper-cast-7079649) Published: 2026-09-04T07:57:57+00:00 Episode link: https://share.transistor.fm/s/e95ac04d Audio file: https://media.transistor.fm/e95ac04d/d1a8d55e.mp3 Processing state: not_requested JSON: https://stenobird.com/v1/public/podcasts/daily-paper-cast-7079649/episodes/rethinking-on-policy-distillation-of-large-language-models-ii-one-training-example Duration seconds: 1332 ## Resource 🤗 Upvotes: 43 | cs.AI, cs.CL Authors: Zixuan Fu, Bingxiang He, Yuxin Zuo, Haohuan Huang, Jinqian Zhang, Ruhang Xiao, Cheng Qian, Qinyu Luo, Huan-ang Gao, Yudong Wang, Zhiyuan Liu, Ning Ding, Chaojun Xiao Title: Rethinking On-Policy Distillation of Large Language Models II: One Training Example Arxiv: http://arxiv.org/abs/2609.04172v1 Abstract: On-policy distillation (OPD) combines student-generated rollouts with dense token-level supervision from a teacher. Existing work has mainly studied its algorithmic behavior, leaving the role of training data unclear. We examine this role at the data-minimal limit by training on a single query. One-shot OPD keeps improving for hundreds of steps and recovers most of full-data OPD's gain across task domains and model families. We explain this result through the states visited during training and the rate at which the student aligns with the teacher. We measure \emph{state coverage}, the fraction of the states full-data OPD visits that a query set's rollouts reach. A single query already reaches \(71.5\%\), most of it within the first 100 steps. Adding semantically distinct queries raises coverage and validation accuracy together, until 16 queries reach \(98.9\%\) and match full-data training. Yet alignment slows at a similar pace whether OPD trains on one query or the whole dataset, and even a fixed set of states takes hundreds of steps to absorb. OPD is therefore data-overfed but algorithm-starved. Its rollouts quickly expose broad supervision, while the student absorbs that supervision increasingly slowly. The state-coverage result extends to multi-teacher OPD, where 16 semantically diverse queries per domain match full-data MOPD. As a further stress test, content-light templates and off-domain WildChat queries also approach the… ## Actions - request_transcript: `POST https://stenobird.com/v1/public/podcasts/daily-paper-cast-7079649/episodes/rethinking-on-policy-distillation-of-large-language-models-ii-one-training-example/transcription-requests` — Idempotently request low-priority transcript generation for this episode. - read_markdown: `GET https://stenobird.com/podcast/daily-paper-cast-7079649/rethinking-on-policy-distillation-of-large-language-models-ii-one-training-example.md` — Read the agent-friendly Markdown representation of this episode resource. A page view does not enqueue transcription. Agents should invoke `request_transcript` explicitly when they need this episode processed. ## Transcript Full transcripts are not published on public pages unless there is a clear rights basis.