# Best AI papers explained Page: https://stenobird.com/podcast/best-ai-papers-explained-7258006 Text version: https://stenobird.com/podcast/best-ai-papers-explained-7258006.md RSS feed: https://anchor.fm/s/1026675f8/podcast/rss Official site: https://podcasters.spotify.com/pod/show/ehwkang Author: Enoch H. Kang Episodes: 789 ## Resource Cut through the noise. We curate and break down the most important AI papers so you don’t have to. ## Machine-readable JSON: https://stenobird.com/v1/public/podcasts/best-ai-papers-explained-7258006 Markdown: https://stenobird.com/podcast/best-ai-papers-explained-7258006.md ## Episodes - [From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning](https://stenobird.com/podcast/best-ai-papers-explained-7258006/from-reasoning-traces-to-reusable-modules-understanding-compositional-generalization-in-language-model-reasoning) — 2026-07-18T16:23:16+00:00 - [Position: Interpretability can be actionable](https://stenobird.com/podcast/best-ai-papers-explained-7258006/position-interpretability-can-be-actionable) — 2026-07-17T16:53:06+00:00 - [High-accuracy sampling for diffusion models and log-concave distributions](https://stenobird.com/podcast/best-ai-papers-explained-7258006/high-accuracy-sampling-for-diffusion-models-and-log-concave-distributions) — 2026-07-17T02:23:28+00:00 - [Causal Inference with Video Features as Treatments](https://stenobird.com/podcast/best-ai-papers-explained-7258006/causal-inference-with-video-features-as-treatments) — 2026-07-15T17:37:45+00:00 - [What Does Thompson Sampling Optimize?](https://stenobird.com/podcast/best-ai-papers-explained-7258006/what-does-thompson-sampling-optimize) — 2026-07-15T00:33:57+00:00 - [Globally Convergent Offline Reinforcement Learning with Smoothed Bellman Residual Minimization](https://stenobird.com/podcast/best-ai-papers-explained-7258006/globally-convergent-offline-reinforcement-learning-with-smoothed-bellman-residual-minimization) — 2026-07-13T17:35:57+00:00 - [LLM-as-a-Verifier: A General-Purpose Verification Framework](https://stenobird.com/podcast/best-ai-papers-explained-7258006/llm-as-a-verifier-a-general-purpose-verification-framework) — 2026-07-10T04:06:20+00:00 - [How Much Do Language Models Memorize?](https://stenobird.com/podcast/best-ai-papers-explained-7258006/how-much-do-language-models-memorize) — 2026-07-09T05:52:54+00:00 - [Position: Uncertainty Quantification in LLMs is Just Unsupervised Clustering](https://stenobird.com/podcast/best-ai-papers-explained-7258006/position-uncertainty-quantification-in-llms-is-just-unsupervised-clustering) — 2026-07-07T17:12:39+00:00 - [Position: Agents Should Invoke External Tools ONLY When Epistemically Necessary](https://stenobird.com/podcast/best-ai-papers-explained-7258006/position-agents-should-invoke-external-tools-only-when-epistemically-necessary) — 2026-07-06T21:50:23+00:00 - [From conversations to mechanisms: aligning advertiser Incentives in ai-powered product recommendations](https://stenobird.com/podcast/best-ai-papers-explained-7258006/from-conversations-to-mechanisms-aligning-advertiser-incentives-in-ai-powered-product-recommendations) — 2026-07-05T20:05:01+00:00 - [Is one layer enough? Training a single transformer layer can match full-parameter RL training](https://stenobird.com/podcast/best-ai-papers-explained-7258006/is-one-layer-enough-training-a-single-transformer-layer-can-match-full-parameter-rl-training) — 2026-07-04T18:08:30+00:00 - [RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training](https://stenobird.com/podcast/best-ai-papers-explained-7258006/rl-excursions-during-pre-training-re-examining-policy-optimization-for-llm-training) — 2026-07-02T22:59:40+00:00 - [Language Generation with Feedback: Queries and Mistakes](https://stenobird.com/podcast/best-ai-papers-explained-7258006/language-generation-with-feedback-queries-and-mistakes) — 2026-07-01T21:10:00+00:00 - [Quantifying Theoretical AI Alignment Guarantees: Receiver-Utility Bounds in Bayesian Persuasion](https://stenobird.com/podcast/best-ai-papers-explained-7258006/quantifying-theoretical-ai-alignment-guarantees-receiver-utility-bounds-in-bayesian-persuasion) — 2026-07-01T03:27:27+00:00 - [SPIRAL: Learning to search and aggregate](https://stenobird.com/podcast/best-ai-papers-explained-7258006/spiral-learning-to-search-and-aggregate) — 2026-06-29T19:54:26+00:00 - [Qwen-AgentWorld: Language World Models for General Agents](https://stenobird.com/podcast/best-ai-papers-explained-7258006/qwen-agentworld-language-world-models-for-general-agents) — 2026-06-27T20:03:10+00:00 - [When Does Trajectory-Level Supervision Permit Efficient Offline Reinforcement Learning?](https://stenobird.com/podcast/best-ai-papers-explained-7258006/when-does-trajectory-level-supervision-permit-efficient-offline-reinforcement-learning) — 2026-06-27T05:11:27+00:00 - [SuperThoughts: Reasoning Tokens in Superposition](https://stenobird.com/podcast/best-ai-papers-explained-7258006/superthoughts-reasoning-tokens-in-superposition) — 2026-06-26T19:42:52+00:00 - [First-Explore PPO : Learning Meta-Exploration with Proximal Policy Optimization](https://stenobird.com/podcast/best-ai-papers-explained-7258006/first-explore-ppo-learning-meta-exploration-with-proximal-policy-optimization) — 2026-06-25T16:09:50+00:00 - [Self-Distillation for Data-Scarce Language Model Pretraining](https://stenobird.com/podcast/best-ai-papers-explained-7258006/self-distillation-for-data-scarce-language-model-pretraining) — 2026-06-24T02:17:55+00:00 - [Meta-Harness for Agent-State Construction](https://stenobird.com/podcast/best-ai-papers-explained-7258006/meta-harness-for-agent-state-construction) — 2026-06-21T15:02:16+00:00 - [ExpRL: Using Reference Solutions as Rewards for LLM Mid-Training](https://stenobird.com/podcast/best-ai-papers-explained-7258006/exprl-using-reference-solutions-as-rewards-for-llm-mid-training) — 2026-06-21T05:28:10+00:00 - [Valid Inference with Synthetic Data via Task Exchangeability](https://stenobird.com/podcast/best-ai-papers-explained-7258006/valid-inference-with-synthetic-data-via-task-exchangeability) — 2026-06-18T22:12:12+00:00 - [GRPO is Secretly a Process Reward Model](https://stenobird.com/podcast/best-ai-papers-explained-7258006/grpo-is-secretly-a-process-reward-model) — 2026-06-17T22:56:07+00:00 ## Actions Episode pages expose an explicit `request_transcript` action. A page view does not automatically enqueue transcription.