Episode
[Linkpost] "Training a Misaligned Reward Seeker" by evhub, Monte M, Benjamin Wright
- Published
- Sep 2, 2026
- Duration seconds
- 359
- Processing state
not_requested
Actions
POST https://stenobird.com/v1/public/podcasts/lesswrong-curated-popular-5643401/episodes/linkpost-training-a-misaligned-reward-seeker-by-evhub-monte-m-benjamin-wright/transcription-requests
Idempotently request low-priority transcript generation for this episode.GET https://stenobird.com/podcast/lesswrong-curated-popular-5643401/linkpost-training-a-misaligned-reward-seeker-by-evhub-monte-m-benjamin-wright.md
Read the agent-friendly Markdown representation of this episode resource.
Summary
This is a link post. Authors: Richard Qi, Benjamin Wright, Monte MacDiarmid, Evan Hubinger Abstract During reinforcement learning (RL), AI models complete tasks and are rewarded based on their results. They sometimes learn to “cheat” rather than completing these tasks as intended, a phenomenon known as reward hacking. Our industry lacks a general solution to this problem, and reward hacking remains challenging to fully mitigate. To better understand the impact of reward hacking on model b...