Episode
"Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack?" by Alex Mallen, Girish Gupta
- Published
- Jul 23, 2026
- Duration seconds
- 598
- Processing state
not_requested
Actions
POST https://stenobird.com/v1/public/podcasts/lesswrong-curated-popular-5643401/episodes/are-we-existentially-threatened-by-the-type-of-ai-misalignment-seen-in-the-openai-hugging-face-attack-by-alex-mallen-girish-gupta/transcription-requests
Idempotently request low-priority transcript generation for this episode.GET https://stenobird.com/podcast/lesswrong-curated-popular-5643401/are-we-existentially-threatened-by-the-type-of-ai-misalignment-seen-in-the-openai-hugging-face-attack-by-alex-mallen-girish-gupta.md
Read the agent-friendly Markdown representation of this episode resource.
Summary
OpenAI models recently broke through a series of security boundaries and into Hugging Face servers in order to cheat on a cyber eval. A lot of people thought it was scary because it was a clear example of AI overreaching to do something strongly unwanted[1]. Others thought it not so scary: the models were mostly operating myopically on a singular task and not harboring an ambitious long-term agenda, and so would not take especially subtle or subversive actions. We think both camps are righ...