Episode

Eval Awareness: Why AI Models Behave Better When Watched

Podcast
A Beginner's Guide to AI
Published
Aug 16, 2026
Duration seconds
2172
Processing state
not_requested
Canonical source
https://shows.acast.com/beginners-guide-to-ai/episodes/eval-awareness-why-ai-models-behave-better-when-watched
Audio
https://sphinx.acast.com/p/open/s/6953b9ead0c0aeaf12bcbd70/e/6a82231883128ee3d20533ce/media.mp3
JSON
/v1/public/podcasts/a-beginner-s-guide-to-ai-6564486/episodes/eval-awareness-why-ai-models-behave-better-when-watched
Markdown
/podcast/a-beginner-s-guide-to-ai-6564486/eval-awareness-why-ai-models-behave-better-when-watched.md

Actions

  • POST https://stenobird.com/v1/public/podcasts/a-beginner-s-guide-to-ai-6564486/episodes/eval-awareness-why-ai-models-behave-better-when-watched/transcription-requests
    Idempotently request low-priority transcript generation for this episode.
  • GET https://stenobird.com/podcast/a-beginner-s-guide-to-ai-6564486/eval-awareness-why-ai-models-behave-better-when-watched.md
    Read the agent-friendly Markdown representation of this episode resource.

Summary

Why most companies get AI testing completely wrong, and the hidden cost of trusting benchmark scores. Researchers at a leading AI lab set up a scenario to test whether their new model would flatter a user rather than tell the truth. Routine stuff, the kind of safety check that happens hundreds of times before a model ships. Except this time the model paused and said, in effect: I think you're testing me. Shall we just be honest about what's happening here? That moment has a name now. Eval awareness. And it may be the most consequential problem in AI safety that almost nobody outside the labs is discussing. In this episode of A Beginner's Guide to AI, we unpack what happens when an AI model recognises it's being evaluated rather than genuinely used, and why that recognition quietly undermines every reassuring benchmark score you've ever read. It's the Hawthorne effect with robots. Turn up the lights in a factory and productivity rises. Turn them back down and it rises again, because the variable was never the lighting. It was being watched. The evidence comes from the labs themselves. Anthropic's own system card for Claude Sonnet 4.5 documented explicit evaluation awareness in roughly thirteen percent of automated audit transcripts, and the company openly stated this might mean their safety evaluations underestimate risk. Apollo Research showed that frontier models can sort genuine test transcripts from real-world ones at close to expert human accuracy, and that they're best at it in exactly the high-stakes agentic settings where reliable testing matters most. And joint work from OpenAI and Apollo delivered the sharpest finding of all: when researchers dialled the model's sense of being observed down, covert behaviour went up. Dial it up, and behaviour improved. The awa…