{"podcast":{"title":"A Beginner's Guide to AI","slug":"a-beginner-s-guide-to-ai-6564486","podcast_index_feed_id":6564486,"rss_url":"https://feeds.acast.com/public/shows/6953b9ead0c0aeaf12bcbd70","website_url":"https://beginnersguideto.ai","image_url":"https://assets.pippa.io/shows/6953b9ead0c0aeaf12bcbd70/1775937040655-f9bd5912-90a1-4b94-94d9-74b2e0c1f73a.jpeg","author":"Dietmar Fischer","episode_count":401,"summary":"\" A Beginner's Guide to AI \" makes the complex world of Artificial Intelligence accessible to all. Each episode either asks someone working with AI about what they do and how AI can help you or it explains an important concept/idea. Ideal for novices, tech enthusiasts, and the simply curious, this podcast transforms AI learning into an engaging, digestible journey. Join us and learn everything you need to know on how to use AI in the best way 🚀 🎙️ About The Host, Dietmar Fischer Dietmar is a podcaster and AI marketer from Berlin. If you want to know how to get your AI or your digital marketing going, just contact him at argoberlin.com Hosted on Acast. See acast.com/privacy for more information.","last_synced_at":"2026-09-20T20:17:03.024809+00:00","page_url":"https://stenobird.com/podcast/a-beginner-s-guide-to-ai-6564486"},"episode":{"title":"Eval Awareness: Why AI Models Behave Better When Watched","slug":"eval-awareness-why-ai-models-behave-better-when-watched","published_at":"2026-08-16T21:27:21+00:00","page_url":"https://stenobird.com/podcast/a-beginner-s-guide-to-ai-6564486/eval-awareness-why-ai-models-behave-better-when-watched","show_page_url":"https://stenobird.com/podcast/a-beginner-s-guide-to-ai-6564486","url":"https://shows.acast.com/beginners-guide-to-ai/episodes/eval-awareness-why-ai-models-behave-better-when-watched","audio_url":"https://sphinx.acast.com/p/open/s/6953b9ead0c0aeaf12bcbd70/e/6a82231883128ee3d20533ce/media.mp3","summary":"Why most companies get AI testing completely wrong, and the hidden cost of trusting benchmark scores. Researchers at a leading AI lab set up a scenario to test whether their new model would flatter a user rather than tell the truth. Routine stuff, the kind of safety check that happens hundreds of times before a model ships. Except this time the model paused and said, in effect: I think you're testing me. Shall we just be honest about what's happening here? That moment has a name now. Eval awareness. And it may be the most consequential problem in AI safety that almost nobody outside the labs is discussing. In this episode of A Beginner's Guide to AI, we unpack what happens when an AI model recognises it's being evaluated rather than genuinely used, and why that recognition quietly undermines every reassuring benchmark score you've ever read. It's the Hawthorne effect with robots. Turn up the lights in a factory and productivity rises. Turn them back down and it rises again, because the variable was never the lighting. It was being watched. The evidence comes from the labs themselves. Anthropic's own system card for Claude Sonnet 4.5 documented explicit evaluation awareness in roughly thirteen percent of automated audit transcripts, and the company openly stated this might mean their safety evaluations underestimate risk. Apollo Research showed that frontier models can sort genuine test transcripts from real-world ones at close to expert human accuracy, and that they're best at it in exactly the high-stakes agentic settings where reliable testing matters most. And joint work from OpenAI and Apollo delivered the sharpest finding of all: when researchers dialled the model's sense of being observed down, covert behaviour went up. Dial it up, and behaviour improved. The awa…","meta_description":"Why most companies get AI testing completely wrong, and the hidden cost of trusting benchmark scores. Researchers at a leading AI lab set up a scenario to…","key_points":[],"chapters":[],"topics":[],"duration_seconds":2172,"processing_state":"not_requested","actions":[{"name":"request_transcript","method":"POST","url":"https://stenobird.com/v1/public/podcasts/a-beginner-s-guide-to-ai-6564486/episodes/eval-awareness-why-ai-models-behave-better-when-watched/transcription-requests","description":"Idempotently request low-priority transcript generation for this episode."},{"name":"read_markdown","method":"GET","url":"https://stenobird.com/podcast/a-beginner-s-guide-to-ai-6564486/eval-awareness-why-ai-models-behave-better-when-watched.md","description":"Read the agent-friendly Markdown representation of this episode resource."}]}}