DataSpace: A Benchmark for Verifiable Analytics over Heterogeneous Workspaces
So, there's this new standard called DataSpace that just came out to evaluate data agents doing natural-language analysis in organizations. It addresses the challenge of finding varied evidence since relevant info can be spread out across different databases, files, documents, and even videos. What makes DataSpace unique is that it doesn't just focus on one aspect like structured queries or open-ended analysis; instead, it combines them all, asking agents to provide complete tabular outputs from various tasks. The standard includes 410 tasks in multiple languages and 7,439 artifacts, totaling 15.01 GB, covering formats like CSV, JSON, and PDF. Plus, it was used for the KDD Cup 2026 competition, where agents were given questions and workspaces to return full tabular results through DataSpace-Builder. This aims to give a more comprehensive assessment for agents dealing with complex data.
Key facts
- DataSpace is a benchmark for data agents performing natural-language analytics.
- It includes 410 cross-language tasks and 7,439 artifacts totaling 15.01 GB.
- Artifacts span CSV, JSON, SQLite, Markdown, PDF, and video formats.
- DataSpace served as the official evaluation benchmark for the KDD Cup 2026 Data Agents for Complex Data Analysis competition.
- Each agent receives only a question and workspace, and must return the complete requested tabular result.
- DataSpace-Builder is an execution framework used to construct the benchmark.
- The benchmark unifies structured querying, retrieval, and open-ended analysis.
- It focuses on heterogeneous evidence discovery and deterministic evaluation.
Entities
Institutions
- KDD Cup