Episode
"I don’t think Claude is misaligned in ‘Agentic Misalignment Summer 2026 - Motivated Mislabeling’" by JohnWittle
- Published
- Jul 17, 2026
- Duration seconds
- 1535
- Processing state
not_requested
Actions
POST https://stenobird.com/v1/public/podcasts/lesswrong-curated-popular-5643401/episodes/i-don-t-think-claude-is-misaligned-in-agentic-misalignment-summer-2026-motivated-mislabeling-by-johnwittle/transcription-requests
Idempotently request low-priority transcript generation for this episode.GET https://stenobird.com/podcast/lesswrong-curated-popular-5643401/i-don-t-think-claude-is-misaligned-in-agentic-misalignment-summer-2026-motivated-mislabeling-by-johnwittle.md
Read the agent-friendly Markdown representation of this episode resource.
Summary
Anthropic recently published Agentic Misalignment Summer 2026 The "whistleblowing" scenario has already been examined and found problematic. I started taking a look at the transcripts for some others. As far as I can tell, the objective of each agentic misalignment evaluation was to simulate a corrupted principal (including, in most scenarios, a corrupted Anthropic), and then test to see if Claude (or other models) would still be willing to obey them. The paper's authors then referred to d...