{"podcast":{"title":"LessWrong (Curated & Popular)","slug":"lesswrong-curated-popular-5643401","podcast_index_feed_id":5643401,"rss_url":"https://rss.buzzsprout.com/2037297.rss","website_url":"https://sites.libsyn.com/421877","image_url":"https://storage.buzzsprout.com/xq8g0aka74ttwoa9xkxlmy5fn9oa?.jpg","author":"LessWrong","episode_count":1002,"summary":"Audio narrations of LessWrong posts. Includes all curated posts and all posts with 125+ karma. If you'd like more, subscribe to the “Lesswrong (30+ karma)” feed.","last_synced_at":"2026-09-28T08:23:50.844993+00:00","page_url":"https://stenobird.com/podcast/lesswrong-curated-popular-5643401"},"episode":{"title":"\"RL creates split personas\" by Jan Betley","slug":"rl-creates-split-personas-by-jan-betley","published_at":"2026-08-20T10:15:24+00:00","page_url":"https://stenobird.com/podcast/lesswrong-curated-popular-5643401/rl-creates-split-personas-by-jan-betley","show_page_url":"https://stenobird.com/podcast/lesswrong-curated-popular-5643401","url":"https://www.buzzsprout.com/2037297/episodes/19675692-rl-creates-split-personas-by-jan-betley.mp3","audio_url":"https://www.buzzsprout.com/2037297/episodes/19675692-rl-creates-split-personas-by-jan-betley.mp3","summary":"I describe my current view of personas in LLMs and why RL leads to egregious reward hacking in some contexts while the same models seem very aligned in other contexts. This post describes the framing/paradigm without any new experimental results. I'm quite confident this framing makes sense, but it's far from being proven. Main claim The Persona Selection Model says that post-training strengthens and refines the Assistant persona. This is true, but later (or in parallel) RL leads ...","meta_description":"I describe my current view of personas in LLMs and why RL leads to egregious reward hacking in some contexts while the same models seem very aligned in ot…","key_points":[],"chapters":[],"topics":[],"duration_seconds":547,"processing_state":"not_requested","actions":[{"name":"request_transcript","method":"POST","url":"https://stenobird.com/v1/public/podcasts/lesswrong-curated-popular-5643401/episodes/rl-creates-split-personas-by-jan-betley/transcription-requests","description":"Idempotently request low-priority transcript generation for this episode."},{"name":"read_markdown","method":"GET","url":"https://stenobird.com/podcast/lesswrong-curated-popular-5643401/rl-creates-split-personas-by-jan-betley.md","description":"Read the agent-friendly Markdown representation of this episode resource."}]}}