Episode
"Empowerment, corrigibility, etc. are simple abstractions (of a messed-up ontology)" by Steven Byrnes
- Published
- Jun 1, 2026
- Duration seconds
- 1864
- Processing state
not_requested
Actions
POST https://stenobird.com/v1/public/podcasts/lesswrong-curated-popular-5643401/episodes/empowerment-corrigibility-etc-are-simple-abstractions-of-a-messed-up-ontology-by-steven-byrnes/transcription-requests
Idempotently request low-priority transcript generation for this episode.GET https://stenobird.com/podcast/lesswrong-curated-popular-5643401/empowerment-corrigibility-etc-are-simple-abstractions-of-a-messed-up-ontology-by-steven-byrnes.md
Read the agent-friendly Markdown representation of this episode resource.
Summary
1.1 Tl;dr Alignment is often conceptualized as AIs helping humans achieve their goals: AIs that increase people's agency and empowerment; AIs that are helpful, corrigible, and/or obedient; AIs that avoid manipulating people. But that last one—manipulation—points to a challenge for all these desiderata: a human's goals are themselves under-determined and manipulable, and it's awfully hard to pin down a principled distinction between changing people's goals in a good way (“providing counsel”...