Episode

[Linkpost] "Frontier models still hack on simple variations of alignment evals from early 2025" by Dean Valentine

Podcast
LessWrong (Curated & Popular)
Published
Sep 8, 2026
Duration seconds
214
Processing state
not_requested
Canonical source
https://www.buzzsprout.com/2037297/episodes/19775271-linkpost-frontier-models-still-hack-on-simple-variations-of-alignment-evals-from-early-2025-by-dean-valentine.mp3
Audio
https://www.buzzsprout.com/2037297/episodes/19775271-linkpost-frontier-models-still-hack-on-simple-variations-of-alignment-evals-from-early-2025-by-dean-valentine.mp3
JSON
/v1/public/podcasts/lesswrong-curated-popular-5643401/episodes/linkpost-frontier-models-still-hack-on-simple-variations-of-alignment-evals-from-early-2025-by-dean-valentine
Markdown
/podcast/lesswrong-curated-popular-5643401/linkpost-frontier-models-still-hack-on-simple-variations-of-alignment-evals-from-early-2025-by-dean-valentine.md

Actions

  • POST https://stenobird.com/v1/public/podcasts/lesswrong-curated-popular-5643401/episodes/linkpost-frontier-models-still-hack-on-simple-variations-of-alignment-evals-from-early-2025-by-dean-valentine/transcription-requests
    Idempotently request low-priority transcript generation for this episode.
  • GET https://stenobird.com/podcast/lesswrong-curated-popular-5643401/linkpost-frontier-models-still-hack-on-simple-variations-of-alignment-evals-from-early-2025-by-dean-valentine.md
    Read the agent-friendly Markdown representation of this episode resource.

Summary

This is a link post. In February 2025, back when o3-mini was the strongest available LLM, Palisade Research publicized a now well-known alignment eval where they asked models to play a game of chess against a chess engine. They found that the new, RLVR'd models cheated on the task by altering the board state about 36% of the time. The experiment received a reasonable amount of circulation, and there were even rumors of skepticism from some lab engineers until they could rerun the evaluation. ...