# [Linkpost] "Frontier models still hack on simple variations of alignment evals from early 2025" by Dean Valentine Page: https://stenobird.com/podcast/lesswrong-curated-popular-5643401/linkpost-frontier-models-still-hack-on-simple-variations-of-alignment-evals-from-early-2025-by-dean-valentine Text version: https://stenobird.com/podcast/lesswrong-curated-popular-5643401/linkpost-frontier-models-still-hack-on-simple-variations-of-alignment-evals-from-early-2025-by-dean-valentine.md Podcast: [LessWrong (Curated & Popular)](https://stenobird.com/podcast/lesswrong-curated-popular-5643401) Published: 2026-09-08T21:45:21+00:00 Episode link: https://www.buzzsprout.com/2037297/episodes/19775271-linkpost-frontier-models-still-hack-on-simple-variations-of-alignment-evals-from-early-2025-by-dean-valentine.mp3 Audio file: https://www.buzzsprout.com/2037297/episodes/19775271-linkpost-frontier-models-still-hack-on-simple-variations-of-alignment-evals-from-early-2025-by-dean-valentine.mp3 Processing state: not_requested JSON: https://stenobird.com/v1/public/podcasts/lesswrong-curated-popular-5643401/episodes/linkpost-frontier-models-still-hack-on-simple-variations-of-alignment-evals-from-early-2025-by-dean-valentine Duration seconds: 214 ## Resource This is a link post. In February 2025, back when o3-mini was the strongest available LLM, Palisade Research publicized a now well-known alignment eval where they asked models to play a game of chess against a chess engine. They found that the new, RLVR'd models cheated on the task by altering the board state about 36% of the time. The experiment received a reasonable amount of circulation, and there were even rumors of skepticism from some lab engineers until they could rerun the evaluation. ... ## Actions - request_transcript: `POST https://stenobird.com/v1/public/podcasts/lesswrong-curated-popular-5643401/episodes/linkpost-frontier-models-still-hack-on-simple-variations-of-alignment-evals-from-early-2025-by-dean-valentine/transcription-requests` — Idempotently request low-priority transcript generation for this episode. - read_markdown: `GET https://stenobird.com/podcast/lesswrong-curated-popular-5643401/linkpost-frontier-models-still-hack-on-simple-variations-of-alignment-evals-from-early-2025-by-dean-valentine.md` — Read the agent-friendly Markdown representation of this episode resource. A page view does not enqueue transcription. Agents should invoke `request_transcript` explicitly when they need this episode processed. ## Transcript Full transcripts are not published on public pages unless there is a clear rights basis.