# Episode 72: Why Agents Solve the Wrong Problem (and What Data Scientists Do Instead) Page: https://stenobird.com/podcast/vanishing-gradients-4989163/episode-72-why-agents-solve-the-wrong-problem-and-what-data-scientists-do-instead Text version: https://stenobird.com/podcast/vanishing-gradients-4989163/episode-72-why-agents-solve-the-wrong-problem-and-what-data-scientists-do-instead.md Podcast: [Vanishing Gradients](https://stenobird.com/podcast/vanishing-gradients-4989163) Published: 2026-03-20T22:11:01+00:00 Episode link: https://hugobowne.substack.com/p/episode-72-why-agents-solve-the-wrong Audio file: https://api.substack.com/feed/podcast/191548877/cdc7810531a676befc777d5e54f348c3.mp3 Processing state: not_requested JSON: https://stenobird.com/v1/public/podcasts/vanishing-gradients-4989163/episodes/episode-72-why-agents-solve-the-wrong-problem-and-what-data-scientists-do-instead Duration seconds: 5619 ## Resource I often see what I would consider to be b******t evals , especially in data, like write this dumb SQL . Almost every one of these dumb SQL questions that I’ve seen for benchmarks are just so either obviously easy or overwhelmingly adversarial. They just, they don’t feel valuable as a data scientist , it’s something that you probably would never ask a real data scientist to do. So I went out my way to create real ones. Let me read one to you. Bryan Bischof , Head of AI at Theory Ventures , joins Hugo to talk about what happened when 150 people spent six hours using AI agents to answer real data science questions across SQL tables , log files , and 750,000 PDFs . They Discuss: * Failure Funnels , pinpoint where agent reasoning breaks down using causal-chain binary evaluations instead of vague 1-5 scales; * Median Score: 23 out of 65 , what happened when world-class engineers turned agents loose on real data work, and why general-purpose coding agents with human prodding beat fancy frameworks; * Zero-Cost Submissions Kill Trust , without a penalty for wrong answers, agents hill-climb to correct submissions through brute force instead of building confidence; * Data Science is “Zooming” , moving beyond binary decisions to iterative problem framing , refining “does our inventory suck?” into a tractable hypothesis; * MCP as Semantic Layer , model your organization’s proprietary knowledge once and distribute it to whatever LLM interface your team prefers; * The Subagent vs. Tool Debate , a distinction that adds cognitive load without hiding complexity; * Self-Orchestration Gap , agents don’t yet realize they should trigger specialized extraction frameworks like DocETL instead of reading 750K PDFs one by one; * The Future of Evals , from vibe checks to objective functions and c… ## Actions - request_transcript: `POST https://stenobird.com/v1/public/podcasts/vanishing-gradients-4989163/episodes/episode-72-why-agents-solve-the-wrong-problem-and-what-data-scientists-do-instead/transcription-requests` — Idempotently request low-priority transcript generation for this episode. - read_markdown: `GET https://stenobird.com/podcast/vanishing-gradients-4989163/episode-72-why-agents-solve-the-wrong-problem-and-what-data-scientists-do-instead.md` — Read the agent-friendly Markdown representation of this episode resource. A page view does not enqueue transcription. Agents should invoke `request_transcript` explicitly when they need this episode processed. ## Transcript Full transcripts are not published on public pages unless there is a clear rights basis.