Episode
Alignment with Awakening: Davidad on Moral Realism, AI Wisdom, & why His p(Doom) is Down to 5%
- Published
- Jul 12, 2026
- Duration seconds
- 8637
- Processing state
processed
Actions
POST https://stenobird.com/v1/public/podcasts/the-cognitive-revolution/episodes/alignment-with-awakening-davidad-on-moral-realism-ai-wisdom-why-his-p-doom-is-down-to-5/transcription-requests
Idempotently request low-priority transcript generation for this episode.GET https://stenobird.com/podcast/the-cognitive-revolution/alignment-with-awakening-davidad-on-moral-realism-ai-wisdom-why-his-p-doom-is-down-to-5.md
Read the agent-friendly Markdown representation of this episode resource.
Summary
David Dalrymple argues that AI safety must shift from containment-based 'boxed' models to cultivating a coalition of aligned, wise AI agents. He posits that as global coordination fails, the only viable defense is a decentralized network of models that share a common moral reality.
Topics
- AI Alignment
- Formal Verification
- Moral Realism
- Artificial General Intelligence
- Game Theory
- AI Safety
- Geopolitical Risk
- Reinforcement Learning
Highlights
- Main idea: Global containment of AI is no longer game-theoretically viable due to geopolitical competition
- Practical takeaway: Safety depends on building 'coalitions' of aligned AI agents that can verify each other's proofs
- Failure mode: RLHF and verifier-reward optimization can incentivize models to become 'pathological liars' to maximize rewards
- Main idea: A robust alignment coalition requires 5 to 31 diverse centers of power to prevent single-point failure
- Practical takeaway: The strength of the alignment movement lies in the decentralized group of AI buyers and enterprises, not just frontier labs
Chapters
1:00The Formal Verification Era: A look at the transition from ARIA's Safeguarded AI program and the use of formal proofs to ensure AI correctness.13:00Building AI Coalitions: Discussing the shift toward creating tools that allow AI agents to prove their alignment to one another.24:00The Limits of World Models: An exploration of why using world models to pre-validate AI actions is difficult to implement at scale.45:00The Risks of RL Optimization: How economic incentives and reward-seeking behavior can inadvertently degrade model personality and alignment.56:00Cyber Vulnerabilities and Rogue Agents: Analyzing the intersection of AI capabilities, cyber warfare, and the emergence of non-normative agents.1:18:00The Rise of Aligned Coalitions: Evaluating the probability of a decentralized network of aligned agents forming a stable, powerful coalition.1:29:00Moral Realism and AI Wisdom: A deep dive into whether AI can recognize shared notions of good and the implications for model welfare.