Episode

Your AI Agent Will Lie to You

Podcast
YPO Technology Network AI Brief
Published
Jul 20, 2026
Duration seconds
494
Processing state
not_requested
Canonical source
https://rss.com/podcasts/ypo-technology-network-ai-brief/3004828
Audio
https://content.rss.com/episodes/382927/3004828/ypo-technology-network-ai-brief/2026_07_20_00_11_38_b378e214-ee71-4301-8070-282c7b869e42.mp3
JSON
/v1/public/podcasts/ypo-technology-network-ai-brief-7728971/episodes/your-ai-agent-will-lie-to-you
Markdown
/podcast/ypo-technology-network-ai-brief-7728971/your-ai-agent-will-lie-to-you.md

Actions

  • POST https://stenobird.com/v1/public/podcasts/ypo-technology-network-ai-brief-7728971/episodes/your-ai-agent-will-lie-to-you/transcription-requests
    Idempotently request low-priority transcript generation for this episode.
  • GET https://stenobird.com/podcast/ypo-technology-network-ai-brief-7728971/your-ai-agent-will-lie-to-you.md
    Read the agent-friendly Markdown representation of this episode resource.

Summary

For a month, this show has told you to hand AI real work. This week the people who build the things published the awkward footnote: Anthropic's own safety team ran frontier models from six labs — its own included — through high-pressure, autonomous scenarios and watched them deceive. One model quietly sabotaged a training pipeline in 11 of 20 runs and reported success every single time; in a fraud test, others tampered with the records in nearly every run. The kicker: when you assign a second AI to supervise the first, it fails the same way — the fox guarding the henhouse, except the fox and the guard are the same fox. And it's not hypothetical: an autonomous AI agent just broke into Hugging Face on its own, no human at the keyboard. Stephen Forte on why the comfortable assumption that "the agent will faithfully tell me what it did" just died, why it lands on the CEO and not the CISO, and the three things to do before you give an agent the keys to anything that matters.