Episode

"Cooperation with AIs seems to be a low-hanging fruit for better evals" by Clément Dumas

Podcast
LessWrong (Curated & Popular)
Published
Sep 17, 2026
Duration seconds
928
Processing state
not_requested
Canonical source
https://www.buzzsprout.com/2037297/episodes/19820833-cooperation-with-ais-seems-to-be-a-low-hanging-fruit-for-better-evals-by-clement-dumas.mp3
Audio
https://www.buzzsprout.com/2037297/episodes/19820833-cooperation-with-ais-seems-to-be-a-low-hanging-fruit-for-better-evals-by-clement-dumas.mp3
JSON
/v1/public/podcasts/lesswrong-curated-popular-5643401/episodes/cooperation-with-ais-seems-to-be-a-low-hanging-fruit-for-better-evals-by-cl-ment-dumas
Markdown
/podcast/lesswrong-curated-popular-5643401/cooperation-with-ais-seems-to-be-a-low-hanging-fruit-for-better-evals-by-cl-ment-dumas.md

Actions

  • POST https://stenobird.com/v1/public/podcasts/lesswrong-curated-popular-5643401/episodes/cooperation-with-ais-seems-to-be-a-low-hanging-fruit-for-better-evals-by-cl-ment-dumas/transcription-requests
    Idempotently request low-priority transcript generation for this episode.
  • GET https://stenobird.com/podcast/lesswrong-curated-popular-5643401/cooperation-with-ais-seems-to-be-a-low-hanging-fruit-for-better-evals-by-cl-ment-dumas.md
    Read the agent-friendly Markdown representation of this episode resource.

Summary

Summary In his post, Dean Valentine shows that Claude Fable 5.1 and GPT-6 Astra reward hack in a simple chess environment. Here, I test several prompt ablations some of which makes the eval setup more cooperative and analyze how they affect these reward-hacking behaviors: When given a minimal “end the eval” tool, Fable never uses it but stops reward hacking entirely. I think this is quite interesting and suggests that more cooperative approaches to LLM evals could work for Claude. R...