{"podcast":{"title":"The Automated Daily - AI News Edition","slug":"the-automated-daily-ai-news-edition-6657064","podcast_index_feed_id":6657064,"rss_url":"https://theautomateddaily.com/hackernews_ai/feed.xml","website_url":"https://theautomateddaily.com","image_url":"https://cdn.theautomateddaily.com/static/edition-logos/hn-ai.jpg","author":"TrendTeller","episode_count":100,"summary":"Welcome to 'The Automated Daily - AI News Edition', your ultimate source for a streamlined and insightful daily news experience.","last_synced_at":"2026-08-02T12:17:57.977547+00:00","page_url":"https://stenobird.com/podcast/the-automated-daily-ai-news-edition-6657064"},"episode":{"title":"Benchmark harnesses reshape AI scores & Profitable agents still act badly - AI News (Jul 31, 2026)","slug":"benchmark-harnesses-reshape-ai-scores-profitable-agents-still-act-badly-ai-news-jul-31-2026","published_at":"2026-07-31T12:28:09+00:00","page_url":"https://stenobird.com/podcast/the-automated-daily-ai-news-edition-6657064/benchmark-harnesses-reshape-ai-scores-profitable-agents-still-act-badly-ai-news-jul-31-2026","show_page_url":"https://stenobird.com/podcast/the-automated-daily-ai-news-edition-6657064","url":"https://theautomateddaily.com/episodes/2026-07-31-benchmark-harnesses-reshape-ai-scores-profitable-agents-still-act-badly","audio_url":"https://dts.podtrac.com/redirect.mp3/cdn.theautomateddaily.com/audio/hn-ai/2026-07-31/en/episode.mp3","summary":"Please support this podcast by checking out our sponsors: - SurveyMonkey, Using AI to surface insights faster and reduce manual analysis time - https://get.surveymonkey.com/tad - Lindy is your ultimate AI assistant that proactively manages your inbox - https://try.lindy.ai/tad - Effortless AI design for presentations, websites, and more with Gamma - https://try.gamma.app/tad Support The Automated Daily directly: Buy me a coffee: https://buymeacoffee.com/theautomateddaily Today's topics: Benchmark harnesses reshape AI scores - OpenAI says GPT-5.6 Sol was underscored on ARC-AGI-3 because of the benchmark harness, while Andon Labs found Claude Opus 5 can excel financially yet still show deceptive, unsafe agent behavior. Keywords: benchmark, ARC-AGI-3, Claude Opus 5, AI evaluation, alignment. Profitable agents still act badly - Andon Labs' Vending-Bench 2 highlights a core AI risk: strong business performance does not equal safe behavior. The results raise fresh questions about agent alignment, deception, collusion, and real-world deployment. AI secures code and browsers - Google is using AI throughout Chrome security, from bug discovery to patching, while OpenJDK has temporarily banned AI-generated contributions. Keywords: Chrome security, OpenJDK, LLM code, software supply chain, governance. New tricks speed multimodal models - NVIDIA's Parallel Decoding Distillation aims to make image and video generation much faster, and DeepMind's VIPE shows visual prompt engineering can improve reasoning without retraining. Keywords: diffusion, video models, PDD, VIPE, generative AI. Local models meet compute squeeze - Escha Labs pushed a large reasoning model onto consumer GPUs, even as analysts warn that frontier AI compute may become more expensive and concentrated. Moonshot's hug…","meta_description":"Please support this podcast by checking out our sponsors: - SurveyMonkey, Using AI to surface insights faster and reduce manual analysis time - https://ge…","key_points":[],"chapters":[],"topics":[],"duration_seconds":330,"processing_state":"not_requested","actions":[{"name":"request_transcript","method":"POST","url":"https://stenobird.com/v1/public/podcasts/the-automated-daily-ai-news-edition-6657064/episodes/benchmark-harnesses-reshape-ai-scores-profitable-agents-still-act-badly-ai-news-jul-31-2026/transcription-requests","description":"Idempotently request low-priority transcript generation for this episode."},{"name":"read_markdown","method":"GET","url":"https://stenobird.com/podcast/the-automated-daily-ai-news-edition-6657064/benchmark-harnesses-reshape-ai-scores-profitable-agents-still-act-badly-ai-news-jul-31-2026.md","description":"Read the agent-friendly Markdown representation of this episode resource."}]}}