{"podcast":{"title":"Vanishing Gradients","slug":"vanishing-gradients-4989163","podcast_index_feed_id":4989163,"rss_url":"https://api.substack.com/feed/podcast/2632531.rss","website_url":"https://hugobowne.substack.com/podcast","image_url":"https://substackcdn.com/feed/podcast/2632531/e8d57d9d781f20857949c2678ef8c9c2.jpg","author":"Hugo Bowne-Anderson","episode_count":77,"summary":"a data podcast with hugo bowne-anderson","last_synced_at":null,"page_url":"https://stenobird.com/podcast/vanishing-gradients-4989163"},"episode":{"title":"Episode 72: Why Agents Solve the Wrong Problem (and What Data Scientists Do Instead)","slug":"episode-72-why-agents-solve-the-wrong-problem-and-what-data-scientists-do-instead","published_at":"2026-03-20T22:11:01+00:00","page_url":"https://stenobird.com/podcast/vanishing-gradients-4989163/episode-72-why-agents-solve-the-wrong-problem-and-what-data-scientists-do-instead","show_page_url":"https://stenobird.com/podcast/vanishing-gradients-4989163","url":"https://hugobowne.substack.com/p/episode-72-why-agents-solve-the-wrong","audio_url":"https://api.substack.com/feed/podcast/191548877/cdc7810531a676befc777d5e54f348c3.mp3","summary":"I often see what I would consider to be b******t evals , especially in data, like write this dumb SQL . Almost every one of these dumb SQL questions that I’ve seen for benchmarks are just so either obviously easy or overwhelmingly adversarial. They just, they don’t feel valuable as a data scientist , it’s something that you probably would never ask a real data scientist to do. So I went out my way to create real ones. Let me read one to you. Bryan Bischof , Head of AI at Theory Ventures , joins Hugo to talk about what happened when 150 people spent six hours using AI agents to answer real data science questions across SQL tables , log files , and 750,000 PDFs . They Discuss: * Failure Funnels , pinpoint where agent reasoning breaks down using causal-chain binary evaluations instead of vague 1-5 scales; * Median Score: 23 out of 65 , what happened when world-class engineers turned agents loose on real data work, and why general-purpose coding agents with human prodding beat fancy frameworks; * Zero-Cost Submissions Kill Trust , without a penalty for wrong answers, agents hill-climb to correct submissions through brute force instead of building confidence; * Data Science is “Zooming” , moving beyond binary decisions to iterative problem framing , refining “does our inventory suck?” into a tractable hypothesis; * MCP as Semantic Layer , model your organization’s proprietary knowledge once and distribute it to whatever LLM interface your team prefers; * The Subagent vs. Tool Debate , a distinction that adds cognitive load without hiding complexity; * Self-Orchestration Gap , agents don’t yet realize they should trigger specialized extraction frameworks like DocETL instead of reading 750K PDFs one by one; * The Future of Evals , from vibe checks to objective functions and c…","meta_description":"I often see what I would consider to be b******t evals , especially in data, like write this dumb SQL . Almost every one of these dumb SQL questions that…","key_points":[],"chapters":[],"topics":[],"duration_seconds":5619,"processing_state":"not_requested","actions":[{"name":"request_transcript","method":"POST","url":"https://stenobird.com/v1/public/podcasts/vanishing-gradients-4989163/episodes/episode-72-why-agents-solve-the-wrong-problem-and-what-data-scientists-do-instead/transcription-requests","description":"Idempotently request low-priority transcript generation for this episode."},{"name":"read_markdown","method":"GET","url":"https://stenobird.com/podcast/vanishing-gradients-4989163/episode-72-why-agents-solve-the-wrong-problem-and-what-data-scientists-do-instead.md","description":"Read the agent-friendly Markdown representation of this episode resource."}]}}