# AgenticDataBench: A Comprehensive Benchmark for Data Agents Page: https://stenobird.com/podcast/daily-paper-cast-7079649/agenticdatabench-a-comprehensive-benchmark-for-data-agents Text version: https://stenobird.com/podcast/daily-paper-cast-7079649/agenticdatabench-a-comprehensive-benchmark-for-data-agents.md Podcast: [Daily Paper Cast](https://stenobird.com/podcast/daily-paper-cast-7079649) Published: 2026-07-04T03:27:13+00:00 Episode link: https://share.transistor.fm/s/cbf201a6 Audio file: https://media.transistor.fm/cbf201a6/9d87f5e8.mp3 Processing state: not_requested JSON: https://stenobird.com/v1/public/podcasts/daily-paper-cast-7079649/episodes/agenticdatabench-a-comprehensive-benchmark-for-data-agents Duration seconds: 1298 ## Resource 🤗 Upvotes: 23 | cs.DB, cs.AI, cs.CL, cs.LG Authors: Zhaoyan Sun, Shan Zhong, Daizhou Wen, Jiaxing Han, Guoliang Li, Ying Yan, Peng Zhang, Yu Su, Xiang Qi, Baolin Sun, Chengyuan Yang, Tao Fang, Huaiyu Ruan Title: AgenticDataBench: A Comprehensive Benchmark for Data Agents Arxiv: http://arxiv.org/abs/2607.01647v1 Abstract: Data science aims to derive actionable insights from heterogeneous raw data, unlocking the value of the massive amounts of data generated in modern society. Automating this process is essential to reducing labor-intensive efforts for data scientists and enabling scalable data-driven applications. Recently, large language model (LLM)-based data agents have emerged as a promising solution to automate data science workflows. However, the field lacks comprehensive benchmarks to rigorously evaluate these agents across diverse scenarios with fine-grained granularity. To address this gap, we propose AgenticDataBench, a comprehensive benchmark featuring realistic tasks spanning diverse domains with fine-grained ground-truth labels. This enables evaluations to capture the diversity and complexity of data science workflows and the detailed performance of agents. First, to cover diverse domains, we collect real datasets and tasks from 15 vertical domains, including 5 real-world B2B use cases from a leading fintech company. Second, to remove redundancy in real-world tasks and generate high-quality tasks for domains lacking real data, we introduce data science skills, recurring data-centric operational patterns, and quantify benchmark coverage by the number of skills included. Representative skills are extracted from large-scale task solutions on Stack Overflow using skill-aligned hierarchical clustering. Third, for real-world business tasks, we select task-solutio… ## Actions - request_transcript: `POST https://stenobird.com/v1/public/podcasts/daily-paper-cast-7079649/episodes/agenticdatabench-a-comprehensive-benchmark-for-data-agents/transcription-requests` — Idempotently request low-priority transcript generation for this episode. - read_markdown: `GET https://stenobird.com/podcast/daily-paper-cast-7079649/agenticdatabench-a-comprehensive-benchmark-for-data-agents.md` — Read the agent-friendly Markdown representation of this episode resource. A page view does not enqueue transcription. Agents should invoke `request_transcript` explicitly when they need this episode processed. ## Transcript Full transcripts are not published on public pages unless there is a clear rights basis.