Frontier LLMs Fail Expert Benchmark on European Executive Tasks
A new benchmark, EuroExec, reveals that frontier large language models (LLMs) fall short of expert judgment on open-ended European executive decision tasks. The study, published on arXiv (2608.04549), involved over 4,000 hours of human expert evaluation. Six frontier LLMs were tested on 413 tasks authored by 47 domain experts, each drawn from real-world cases. Responses were manually assessed using a multi-attribute rubric, item-specific checklists, and preference rankings, yielding an aggregate 'Solve Rate'. The strongest model solved only 56.9% of tasks, while expert-written reference answers achieved near-ceiling performance and were preferred over every model response in 74% of direct rankings. This places frontier generative systems well below professional standards for such complex, open-ended challenges.
Key facts
- EuroExec is a new benchmark for European executive decision tasks.
- The study involved over 4,000 human expert hours.
- Six frontier LLMs were evaluated.
- The benchmark consists of 413 tasks authored by 47 domain experts.
- Tasks are open-ended and long-form, based on real cases.
- Responses were evaluated using a multi-attribute rubric and item-specific checklists.
- The strongest model solved only 56.9% of tasks.
- Expert-written reference answers were preferred over model responses in 74% of direct rankings.
Entities
Institutions
- arXiv