{"podcast":{"title":"LessWrong (Curated & Popular)","slug":"lesswrong-curated-popular-5643401","podcast_index_feed_id":5643401,"rss_url":"https://rss.buzzsprout.com/2037297.rss","website_url":"https://sites.libsyn.com/421877","image_url":"https://storage.buzzsprout.com/xq8g0aka74ttwoa9xkxlmy5fn9oa?.jpg","author":"LessWrong","episode_count":1002,"summary":"Audio narrations of LessWrong posts. Includes all curated posts and all posts with 125+ karma. If you'd like more, subscribe to the “Lesswrong (30+ karma)” feed.","last_synced_at":"2026-09-28T08:23:50.844993+00:00","page_url":"https://stenobird.com/podcast/lesswrong-curated-popular-5643401"},"episode":{"title":"[Linkpost] \"Frontier models still hack on simple variations of alignment evals from early 2025\" by Dean Valentine","slug":"linkpost-frontier-models-still-hack-on-simple-variations-of-alignment-evals-from-early-2025-by-dean-valentine","published_at":"2026-09-08T21:45:21+00:00","page_url":"https://stenobird.com/podcast/lesswrong-curated-popular-5643401/linkpost-frontier-models-still-hack-on-simple-variations-of-alignment-evals-from-early-2025-by-dean-valentine","show_page_url":"https://stenobird.com/podcast/lesswrong-curated-popular-5643401","url":"https://www.buzzsprout.com/2037297/episodes/19775271-linkpost-frontier-models-still-hack-on-simple-variations-of-alignment-evals-from-early-2025-by-dean-valentine.mp3","audio_url":"https://www.buzzsprout.com/2037297/episodes/19775271-linkpost-frontier-models-still-hack-on-simple-variations-of-alignment-evals-from-early-2025-by-dean-valentine.mp3","summary":"This is a link post. In February 2025, back when o3-mini was the strongest available LLM, Palisade Research publicized a now well-known alignment eval where they asked models to play a game of chess against a chess engine. They found that the new, RLVR'd models cheated on the task by altering the board state about 36% of the time. The experiment received a reasonable amount of circulation, and there were even rumors of skepticism from some lab engineers until they could rerun the evaluation. ...","meta_description":"This is a link post. In February 2025, back when o3-mini was the strongest available LLM, Palisade Research publicized a now well-known alignment eval whe…","key_points":[],"chapters":[],"topics":[],"duration_seconds":214,"processing_state":"not_requested","actions":[{"name":"request_transcript","method":"POST","url":"https://stenobird.com/v1/public/podcasts/lesswrong-curated-popular-5643401/episodes/linkpost-frontier-models-still-hack-on-simple-variations-of-alignment-evals-from-early-2025-by-dean-valentine/transcription-requests","description":"Idempotently request low-priority transcript generation for this episode."},{"name":"read_markdown","method":"GET","url":"https://stenobird.com/podcast/lesswrong-curated-popular-5643401/linkpost-frontier-models-still-hack-on-simple-variations-of-alignment-evals-from-early-2025-by-dean-valentine.md","description":"Read the agent-friendly Markdown representation of this episode resource."}]}}