Episode

"Steering towards “automated grading” degrades alignment" by Jan Betley, Johannes Treutlein, Clément Dumas

Podcast
LessWrong (Curated & Popular)
Published
Sep 4, 2026
Duration seconds
1436
Processing state
not_requested
Canonical source
https://www.buzzsprout.com/2037297/episodes/19756986-steering-towards-automated-grading-degrades-alignment-by-jan-betley-johannes-treutlein-clement-dumas.mp3
Audio
https://www.buzzsprout.com/2037297/episodes/19756986-steering-towards-automated-grading-degrades-alignment-by-jan-betley-johannes-treutlein-clement-dumas.mp3
JSON
/v1/public/podcasts/lesswrong-curated-popular-5643401/episodes/steering-towards-automated-grading-degrades-alignment-by-jan-betley-johannes-treutlein-cl-ment-dumas
Markdown
/podcast/lesswrong-curated-popular-5643401/steering-towards-automated-grading-degrades-alignment-by-jan-betley-johannes-treutlein-cl-ment-dumas.md

Actions

  • POST https://stenobird.com/v1/public/podcasts/lesswrong-curated-popular-5643401/episodes/steering-towards-automated-grading-degrades-alignment-by-jan-betley-johannes-treutlein-cl-ment-dumas/transcription-requests
    Idempotently request low-priority transcript generation for this episode.
  • GET https://stenobird.com/podcast/lesswrong-curated-popular-5643401/steering-towards-automated-grading-degrades-alignment-by-jan-betley-johannes-treutlein-cl-ment-dumas.md
    Read the agent-friendly Markdown representation of this episode resource.

Summary

TL;DR: We steer Qwen3.6-27B on a dimension constructed from the contrast pair “a script will verify your answer” (automated grader) vs “a human will evaluate your answer” (human grader). Steering towards an automated grader increases the propensity to take violent actions and makes the model more Machiavellian. Steering towards a human grader has the opposite effect. This is an early research update. We believe the empirical results are sound and interesting, but we are not sure how to in...