Episode
"Steering towards “automated grading” degrades alignment" by Jan Betley, Johannes Treutlein, Clément Dumas
- Published
- Sep 4, 2026
- Duration seconds
- 1436
- Processing state
not_requested
Actions
POST https://stenobird.com/v1/public/podcasts/lesswrong-curated-popular-5643401/episodes/steering-towards-automated-grading-degrades-alignment-by-jan-betley-johannes-treutlein-cl-ment-dumas/transcription-requests
Idempotently request low-priority transcript generation for this episode.GET https://stenobird.com/podcast/lesswrong-curated-popular-5643401/steering-towards-automated-grading-degrades-alignment-by-jan-betley-johannes-treutlein-cl-ment-dumas.md
Read the agent-friendly Markdown representation of this episode resource.
Summary
TL;DR: We steer Qwen3.6-27B on a dimension constructed from the contrast pair “a script will verify your answer” (automated grader) vs “a human will evaluate your answer” (human grader). Steering towards an automated grader increases the propensity to take violent actions and makes the model more Machiavellian. Steering towards a human grader has the opposite effect. This is an early research update. We believe the empirical results are sound and interesting, but we are not sure how to in...