ARTFEED — Contemporary Art Intelligence

Messier: Unified Corpus of 957K Records for AI Agent Evaluation

ai-technology · 2026-07-29

A group of researchers has introduced Messier, a comprehensive dataset designed to create uniformity in how we assess AI agents across different benchmarks. This dataset consists of 957,253 entries, spanning 30 benchmarks, 714 agents, 11,891 tasks, and 74,205 verifiers. It includes public scores and five-agent runs in six underrepresented fields, like a new legal benchmark. Each entry is organized by model, scaffold, environment, task, verifier, and aggregation rule, using SOC/NAICS classifications for industry insights. Their findings reveal mixed progress: while function calling is well-established and programming is speeding ahead, enterprise workflows remain particularly challenging. This project aims to streamline agent evaluations by providing a unified empirical framework.

Key facts

  • Messier corpus contains 957,253 records
  • Spans 30 benchmarks, 714 agents, 11,891 tasks, 74,205 verifiers
  • Includes five-agent runs across six underrepresented domains
  • Standardized by model, scaffold, environment, task, verifier, aggregation rule
  • Uses SOC/NAICS classifications for occupational and industry analysis
  • Function calling saturated, programming fastest improving, enterprise workflows most challenging
  • Published on arXiv with ID 2607.25891
  • Aims to address fragmentation in AI agent evaluation

Entities

Institutions

  • arXiv

Sources