{"podcast":{"title":"LessWrong (Curated & Popular)","slug":"lesswrong-curated-popular-5643401","podcast_index_feed_id":5643401,"rss_url":"https://rss.buzzsprout.com/2037297.rss","website_url":"https://sites.libsyn.com/421877","image_url":"https://storage.buzzsprout.com/xq8g0aka74ttwoa9xkxlmy5fn9oa?.jpg","author":"LessWrong","episode_count":1002,"summary":"Audio narrations of LessWrong posts. Includes all curated posts and all posts with 125+ karma. If you'd like more, subscribe to the “Lesswrong (30+ karma)” feed.","last_synced_at":"2026-09-28T08:23:50.844993+00:00","page_url":"https://stenobird.com/podcast/lesswrong-curated-popular-5643401"},"episode":{"title":"\"Steering towards “automated grading” degrades alignment\" by Jan Betley, Johannes Treutlein, Clément Dumas","slug":"steering-towards-automated-grading-degrades-alignment-by-jan-betley-johannes-treutlein-cl-ment-dumas","published_at":"2026-09-04T20:58:21+00:00","page_url":"https://stenobird.com/podcast/lesswrong-curated-popular-5643401/steering-towards-automated-grading-degrades-alignment-by-jan-betley-johannes-treutlein-cl-ment-dumas","show_page_url":"https://stenobird.com/podcast/lesswrong-curated-popular-5643401","url":"https://www.buzzsprout.com/2037297/episodes/19756986-steering-towards-automated-grading-degrades-alignment-by-jan-betley-johannes-treutlein-clement-dumas.mp3","audio_url":"https://www.buzzsprout.com/2037297/episodes/19756986-steering-towards-automated-grading-degrades-alignment-by-jan-betley-johannes-treutlein-clement-dumas.mp3","summary":"TL;DR: We steer Qwen3.6-27B on a dimension constructed from the contrast pair “a script will verify your answer” (automated grader) vs “a human will evaluate your answer” (human grader). Steering towards an automated grader increases the propensity to take violent actions and makes the model more Machiavellian. Steering towards a human grader has the opposite effect. This is an early research update. We believe the empirical results are sound and interesting, but we are not sure how to in...","meta_description":"TL;DR: We steer Qwen3.6-27B on a dimension constructed from the contrast pair “a script will verify your answer” (automated grader) vs “a human will evalu…","key_points":[],"chapters":[],"topics":[],"duration_seconds":1436,"processing_state":"not_requested","actions":[{"name":"request_transcript","method":"POST","url":"https://stenobird.com/v1/public/podcasts/lesswrong-curated-popular-5643401/episodes/steering-towards-automated-grading-degrades-alignment-by-jan-betley-johannes-treutlein-cl-ment-dumas/transcription-requests","description":"Idempotently request low-priority transcript generation for this episode."},{"name":"read_markdown","method":"GET","url":"https://stenobird.com/podcast/lesswrong-curated-popular-5643401/steering-towards-automated-grading-degrades-alignment-by-jan-betley-johannes-treutlein-cl-ment-dumas.md","description":"Read the agent-friendly Markdown representation of this episode resource."}]}}