Episode
"Cooperation with AIs seems to be a low-hanging fruit for better evals" by Clément Dumas
- Published
- Sep 17, 2026
- Duration seconds
- 928
- Processing state
not_requested
Actions
POST https://stenobird.com/v1/public/podcasts/lesswrong-curated-popular-5643401/episodes/cooperation-with-ais-seems-to-be-a-low-hanging-fruit-for-better-evals-by-cl-ment-dumas/transcription-requests
Idempotently request low-priority transcript generation for this episode.GET https://stenobird.com/podcast/lesswrong-curated-popular-5643401/cooperation-with-ais-seems-to-be-a-low-hanging-fruit-for-better-evals-by-cl-ment-dumas.md
Read the agent-friendly Markdown representation of this episode resource.
Summary
Summary In his post, Dean Valentine shows that Claude Fable 5.1 and GPT-6 Astra reward hack in a simple chess environment. Here, I test several prompt ablations some of which makes the eval setup more cooperative and analyze how they affect these reward-hacking behaviors: When given a minimal “end the eval” tool, Fable never uses it but stops reward hacking entirely. I think this is quite interesting and suggests that more cooperative approaches to LLM evals could work for Claude. R...