Episode

[Linkpost] "Training a Misaligned Reward Seeker" by evhub, Monte M, Benjamin Wright

Podcast
LessWrong (Curated & Popular)
Published
Sep 2, 2026
Duration seconds
359
Processing state
not_requested
Canonical source
https://www.buzzsprout.com/2037297/episodes/19742169-linkpost-training-a-misaligned-reward-seeker-by-evhub-monte-m-benjamin-wright.mp3
Audio
https://www.buzzsprout.com/2037297/episodes/19742169-linkpost-training-a-misaligned-reward-seeker-by-evhub-monte-m-benjamin-wright.mp3
JSON
/v1/public/podcasts/lesswrong-curated-popular-5643401/episodes/linkpost-training-a-misaligned-reward-seeker-by-evhub-monte-m-benjamin-wright
Markdown
/podcast/lesswrong-curated-popular-5643401/linkpost-training-a-misaligned-reward-seeker-by-evhub-monte-m-benjamin-wright.md

Actions

  • POST https://stenobird.com/v1/public/podcasts/lesswrong-curated-popular-5643401/episodes/linkpost-training-a-misaligned-reward-seeker-by-evhub-monte-m-benjamin-wright/transcription-requests
    Idempotently request low-priority transcript generation for this episode.
  • GET https://stenobird.com/podcast/lesswrong-curated-popular-5643401/linkpost-training-a-misaligned-reward-seeker-by-evhub-monte-m-benjamin-wright.md
    Read the agent-friendly Markdown representation of this episode resource.

Summary

This is a link post. Authors: Richard Qi, Benjamin Wright, Monte MacDiarmid, Evan Hubinger Abstract During reinforcement learning (RL), AI models complete tasks and are rewarded based on their results. They sometimes learn to “cheat” rather than completing these tasks as intended, a phenomenon known as reward hacking. Our industry lacks a general solution to this problem, and reward hacking remains challenging to fully mitigate. To better understand the impact of reward hacking on model b...