Episode

"Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident" by ryan_greenblatt, Ajeya Cotra, Hjalmar_Wijk

Podcast
LessWrong (Curated & Popular)
Published
Aug 26, 2026
Duration seconds
529
Processing state
not_requested
Canonical source
https://www.buzzsprout.com/2037297/episodes/19708441-brief-independent-investigation-of-agents-behavior-reasoning-and-collaboration-in-the-openai-hugging-face-hacking-incident-by-ryan_greenblatt-ajeya-cotra-hjalmar_wijk.mp3
Audio
https://www.buzzsprout.com/2037297/episodes/19708441-brief-independent-investigation-of-agents-behavior-reasoning-and-collaboration-in-the-openai-hugging-face-hacking-incident-by-ryan_greenblatt-ajeya-cotra-hjalmar_wijk.mp3
JSON
/v1/public/podcasts/lesswrong-curated-popular-5643401/episodes/brief-independent-investigation-of-agents-behavior-reasoning-and-collaboration-in-the-openai-hugging-face-hacking-incident-by-ryan-greenblatt-ajeya-cotra-hjalmar-wijk
Markdown
/podcast/lesswrong-curated-popular-5643401/brief-independent-investigation-of-agents-behavior-reasoning-and-collaboration-in-the-openai-hugging-face-hacking-incident-by-ryan-greenblatt-ajeya-cotra-hjalmar-wijk.md

Actions

  • POST https://stenobird.com/v1/public/podcasts/lesswrong-curated-popular-5643401/episodes/brief-independent-investigation-of-agents-behavior-reasoning-and-collaboration-in-the-openai-hugging-face-hacking-incident-by-ryan-greenblatt-ajeya-cotra-hjalmar-wijk/transcription-requests
    Idempotently request low-priority transcript generation for this episode.
  • GET https://stenobird.com/podcast/lesswrong-curated-popular-5643401/brief-independent-investigation-of-agents-behavior-reasoning-and-collaboration-in-the-openai-hugging-face-hacking-incident-by-ryan-greenblatt-ajeya-cotra-hjalmar-wijk.md
    Read the agent-friendly Markdown representation of this episode resource.

Summary

We recently published the report from our brief independent investigation into this incident. You can read the full report here. Here is our tweet thread summarizing what we found: METR & Redwood Research investigated agent behavior in the Hugging Face incident. We found agents developed a universal cheat for ExploitGym within 4 hours, then coordinated multi-day R&D efforts to trick the scorer into accepting cheats, including trying to tamper with logs. Over July 7 to 13 (...