# ExpRL: Using Reference Solutions as Rewards for LLM Mid-Training Page: https://stenobird.com/podcast/best-ai-papers-explained-7258006/exprl-using-reference-solutions-as-rewards-for-llm-mid-training Text version: https://stenobird.com/podcast/best-ai-papers-explained-7258006/exprl-using-reference-solutions-as-rewards-for-llm-mid-training.md Podcast: [Best AI papers explained](https://stenobird.com/podcast/best-ai-papers-explained-7258006) Published: 2026-06-21T05:28:10+00:00 Episode link: https://podcasters.spotify.com/pod/show/ehwkang/episodes/ExpRL-Using-Reference-Solutions-as-Rewards-for-LLM-Mid-Training-e3l2mpb Audio file: https://anchor.fm/s/1026675f8/podcast/play/121772267/https%3A%2F%2Fd3ctxlq1ktw2nl.cloudfront.net%2Fstaging%2F2026-5-21%2F338cc5a8-0973-7322-1365-c4436ce1b824.m4a Processing state: not_requested JSON: https://stenobird.com/v1/public/podcasts/best-ai-papers-explained-7258006/episodes/exprl-using-reference-solutions-as-rewards-for-llm-mid-training Duration seconds: 1263 ## Resource Exploratory RL (ExpRL) is an automated mid-training method designed to enhance the reasoning capabilities of large language models before they undergo standard reinforcement learning. While traditional reinforcement learning often struggles with sparse rewards on difficult problems, ExpRL uses human-written reference solutions as reward scaffolds to provide dense, informative feedback on partial progress. This approach employs an LLM judge to evaluate on-policy reasoning traces against specific rubrics, assigning rewards at both the outcome and process levels to reinforce productive intermediate steps. By shifting probability mass toward successful solution strategies, the method significantly improves pass@k performance and broadens the model’s coverage of complex reasoning paths. Experimental results demonstrate that ExpRL creates a superior initialization for subsequent training, outperforming supervised fine-tuning and standard distillation across challenging math and science benchmarks. Ultimately, this technique fosters sophisticated behaviors like self-correction and backtracking, which are essential for solving high-level reasoning tasks. ## Actions - request_transcript: `POST https://stenobird.com/v1/public/podcasts/best-ai-papers-explained-7258006/episodes/exprl-using-reference-solutions-as-rewards-for-llm-mid-training/transcription-requests` — Idempotently request low-priority transcript generation for this episode. - read_markdown: `GET https://stenobird.com/podcast/best-ai-papers-explained-7258006/exprl-using-reference-solutions-as-rewards-for-llm-mid-training.md` — Read the agent-friendly Markdown representation of this episode resource. A page view does not enqueue transcription. Agents should invoke `request_transcript` explicitly when they need this episode processed. ## Transcript Full transcripts are not published on public pages unless there is a clear rights basis.