{"podcast":{"title":"Best AI papers explained","slug":"best-ai-papers-explained-7258006","podcast_index_feed_id":7258006,"rss_url":"https://anchor.fm/s/1026675f8/podcast/rss","website_url":"https://podcasters.spotify.com/pod/show/ehwkang","image_url":"https://d3t3ozftmdmh3i.cloudfront.net/staging/podcast_uploaded_nologo/43252366/43252366-1744500070152-e62b760188d8.jpg","author":"Enoch H. Kang","episode_count":789,"summary":"Cut through the noise. We curate and break down the most important AI papers so you don’t have to.","last_synced_at":"2026-07-19T16:17:08.576018+00:00","page_url":"https://stenobird.com/podcast/best-ai-papers-explained-7258006"},"episode":{"title":"When Does Trajectory-Level Supervision Permit Efficient Offline Reinforcement Learning?","slug":"when-does-trajectory-level-supervision-permit-efficient-offline-reinforcement-learning","published_at":"2026-06-27T05:11:27+00:00","page_url":"https://stenobird.com/podcast/best-ai-papers-explained-7258006/when-does-trajectory-level-supervision-permit-efficient-offline-reinforcement-learning","show_page_url":"https://stenobird.com/podcast/best-ai-papers-explained-7258006","url":"https://podcasters.spotify.com/pod/show/ehwkang/episodes/When-Does-Trajectory-Level-Supervision-Permit-Efficient-Offline-Reinforcement-Learning-e3lb91k","audio_url":"https://anchor.fm/s/1026675f8/podcast/play/122053108/https%3A%2F%2Fd3ctxlq1ktw2nl.cloudfront.net%2Fstaging%2F2026-5-27%2F99cded2d-7fc3-eb3d-aa06-8cede015dc9c.m4a","summary":"This paper discusses a statistical framework for offline reinforcement learning using trajectory-level supervision, where only final outcomes or preferences are observed rather than step-by-step rewards. The authors introduce OPAC, a pessimistic actor-critic algorithm designed to learn from these aggregated signals by estimating latent rewards and applying pessimism to account for distribution shifts. Their analysis establishes that moving from process-level to outcome-level feedback incurs a quantifiable statistical cost, specifically an additional horizon factor in sample complexity. The research also explores generalized RL objectives, proving that non-linear outcomes like &quot;all-success&quot; criteria can lead to exponentially difficult learning problems. To address this, they identify specific structural coefficients, $\\kappa_\\mu(\\sigma)$ and $\\chi_\\mu(\\sigma)$, which determine when efficient learning remains possible. Ultimately, the paper provides a theoretical boundary for when sparse, trajectory-based data can successfully guide sequential decision-making.","meta_description":"This paper discusses a statistical framework for offline reinforcement learning using trajectory-level supervision, where only final outcomes or preferenc…","key_points":[],"chapters":[],"topics":[],"duration_seconds":1136,"processing_state":"not_requested","actions":[{"name":"request_transcript","method":"POST","url":"https://stenobird.com/v1/public/podcasts/best-ai-papers-explained-7258006/episodes/when-does-trajectory-level-supervision-permit-efficient-offline-reinforcement-learning/transcription-requests","description":"Idempotently request low-priority transcript generation for this episode."},{"name":"read_markdown","method":"GET","url":"https://stenobird.com/podcast/best-ai-papers-explained-7258006/when-does-trajectory-level-supervision-permit-efficient-offline-reinforcement-learning.md","description":"Read the agent-friendly Markdown representation of this episode resource."}]}}