{"podcast":{"title":"Best AI papers explained","slug":"best-ai-papers-explained-7258006","podcast_index_feed_id":7258006,"rss_url":"https://anchor.fm/s/1026675f8/podcast/rss","website_url":"https://podcasters.spotify.com/pod/show/ehwkang","image_url":"https://d3t3ozftmdmh3i.cloudfront.net/staging/podcast_uploaded_nologo/43252366/43252366-1744500070152-e62b760188d8.jpg","author":"Enoch H. Kang","episode_count":789,"summary":"Cut through the noise. We curate and break down the most important AI papers so you don’t have to.","last_synced_at":"2026-07-19T16:17:08.576018+00:00","page_url":"https://stenobird.com/podcast/best-ai-papers-explained-7258006"},"episode":{"title":"RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training","slug":"rl-excursions-during-pre-training-re-examining-policy-optimization-for-llm-training","published_at":"2026-07-02T22:59:40+00:00","page_url":"https://stenobird.com/podcast/best-ai-papers-explained-7258006/rl-excursions-during-pre-training-re-examining-policy-optimization-for-llm-training","show_page_url":"https://stenobird.com/podcast/best-ai-papers-explained-7258006","url":"https://podcasters.spotify.com/pod/show/ehwkang/episodes/RL-Excursions-during-Pre-Training-Re-examining-Policy-Optimization-for-LLM-training-e3lja5t","audio_url":"https://anchor.fm/s/1026675f8/podcast/play/122316413/https%3A%2F%2Fd3ctxlq1ktw2nl.cloudfront.net%2Fstaging%2F2026-6-2%2Fd5fccd4c-dd68-eff0-ee81-59d50b263cb5.m4a","summary":"This research investigates the effectiveness of integrating reinforcement learning (RL) earlier in the large language model training pipeline rather than treating it solely as a final post-training step. The authors demonstrate that RL is effective remarkably early, often matching the performance of standard sequential pipelines after only a small fraction of pre-training is complete. Unlike supervised fine-tuning (SFT), which tends to degrade a model's general capabilities and narrow its output, direct RL preserves general skills and expands the diversity of reasoning paths. The study also identifies that targeted data composition is more critical for RL success than simply increasing model size. Finally, the researchers propose a parallel averaging method that combines RL and SFT updates to achieve superior results across all training stages. Together, these findings suggest that the current standard of isolating RL to the end of training is an unnecessary design choice that limits model potential.","meta_description":"This research investigates the effectiveness of integrating reinforcement learning (RL) earlier in the large language model training pipeline rather than…","key_points":[],"chapters":[],"topics":[],"duration_seconds":1307,"processing_state":"not_requested","actions":[{"name":"request_transcript","method":"POST","url":"https://stenobird.com/v1/public/podcasts/best-ai-papers-explained-7258006/episodes/rl-excursions-during-pre-training-re-examining-policy-optimization-for-llm-training/transcription-requests","description":"Idempotently request low-priority transcript generation for this episode."},{"name":"read_markdown","method":"GET","url":"https://stenobird.com/podcast/best-ai-papers-explained-7258006/rl-excursions-during-pre-training-re-examining-policy-optimization-for-llm-training.md","description":"Read the agent-friendly Markdown representation of this episode resource."}]}}