ARTFEED — Contemporary Art Intelligence

DSAgentBench: Benchmarking AI Agents for Real-World Data Science Workflows

ai-technology · 2026-08-13

The introduction of a new benchmark, DSAgentBench, aims to assess the ability of AI agents to fully automate end-to-end data-science workflows in authentic computing environments. This benchmark, explained in a paper available on arXiv (arXiv:2608.10366), fills a void in current benchmarks that do not involve real-computer interactions and overlook the multi-stage, multi-tool aspects of data science. DSAgentBench features 275 varied tasks that encompass the entire data-science life cycle, including data wrangling, exploration, modeling, visualization, and validation. Agents must effectively utilize tools like notebooks, IDEs, terminals, browsers, and databases in real operating settings. Each task necessitates informed decision-making based on intermediate outputs and includes a deterministic evaluation, reflecting the intricate coordination required in practical data science.

Key facts

  • DSAgentBench is the first benchmark to evaluate agents on full data-science workflows in real computer environments.
  • It contains 275 diverse tasks covering the entire data-science life-cycle.
  • Tasks require coordination of tools like notebooks, IDEs, terminals, browsers, and databases.
  • The benchmark addresses the lack of real-computer interaction in existing benchmarks.
  • Each task includes deterministic evaluation.
  • The paper is available on arXiv with identifier 2608.10366.
  • The benchmark reflects the multi-stage, multi-tool nature of data-science practice.

Entities

Institutions

  • arXiv

Sources