# Pass the Baton: Trajectory-Relayed On-Policy Distillation Page: https://stenobird.com/podcast/daily-paper-cast-7079649/pass-the-baton-trajectory-relayed-on-policy-distillation Text version: https://stenobird.com/podcast/daily-paper-cast-7079649/pass-the-baton-trajectory-relayed-on-policy-distillation.md Podcast: [Daily Paper Cast](https://stenobird.com/podcast/daily-paper-cast-7079649) Published: 2026-07-30T03:18:12+00:00 Episode link: https://share.transistor.fm/s/2c945e76 Audio file: https://media.transistor.fm/2c945e76/7a796624.mp3 Processing state: not_requested JSON: https://stenobird.com/v1/public/podcasts/daily-paper-cast-7079649/episodes/pass-the-baton-trajectory-relayed-on-policy-distillation Duration seconds: 1219 ## Resource 🤗 Upvotes: 24 | cs.CL, cs.AI Authors: Haolei Xu, Xiaowen Xu, Haiwen Hong, Zixuan Ni, Hongxing Li, Yiwen Qiu, Weiming Lu, Yongliang Shen Title: Pass the Baton: Trajectory-Relayed On-Policy Distillation Arxiv: http://arxiv.org/abs/2607.26057v1 Abstract: On-policy distillation (OPD) grounds token-level supervision in the student's own trajectory, yet suffers from prefix failure: once the student commits to a wrong reasoning direction, all subsequent generation builds on this deviation, producing misdirected continuations that elicit unreliable supervision and waste compute. We identify a teacher-student continuation asymmetry on failed prefixes, where the teacher tends to redirect while the student continues along the original direction, and convert it into a label-free handoff trigger in Relay On-Policy Distillation (Relay-OPD). During training, Relay-OPD constructs relay trajectories by letting the teacher briefly take over at detected trigger points to produce a teacher leg, after which the student resumes and is optimized on the resulting trajectory. A limited relay budget concentrates intervention on critical early positions while limiting departure from the student policy. With a Qwen3-4B-Instruct-2507 teacher and Qwen3-0.6B/1.7B-Non-Thinking students on eight mathematical reasoning benchmarks, Relay-OPD achieves the best or second-best results on every benchmark, outperforming standard OPD by +5.73% and the strongest baseline FastOPD by +1.49% on average for 1.7B, with consistent gains at 0.6B. Training trajectory length is reduced by over 50%. ## Actions - request_transcript: `POST https://stenobird.com/v1/public/podcasts/daily-paper-cast-7079649/episodes/pass-the-baton-trajectory-relayed-on-policy-distillation/transcription-requests` — Idempotently request low-priority transcript generation for this episode. - read_markdown: `GET https://stenobird.com/podcast/daily-paper-cast-7079649/pass-the-baton-trajectory-relayed-on-policy-distillation.md` — Read the agent-friendly Markdown representation of this episode resource. A page view does not enqueue transcription. Agents should invoke `request_transcript` explicitly when they need this episode processed. ## Transcript Full transcripts are not published on public pages unless there is a clear rights basis.