DashArena: New Benchmark for Interactive Analytic Dashboard Generation
Researchers have introduced a new benchmark called DashArena, aimed at evaluating large language models (LLMs) in their ability to create interactive analytic dashboards. The details are shared in a paper available on arXiv (ID: 2608.10567). This benchmark tackles the challenge of assessing dashboard generation, which is often open-ended. DashArena stands out as the first of its kind for task-grounded and open-ended dashboard creation. It requires models to produce a dashboard along with a replayable interaction trajectory. A browser executor then replays this path, providing visual proof of the analytic process. A vision-language model (VLM) judge assesses the outputs, and a leaderboard is created using Bradley–Terry aggregation. The researchers also developed the DashJudge-8B model based on the judge. You can read the paper at https://arxiv.org/abs/2608.10567.
Key facts
- DashArena is introduced as a benchmark for interactive analytic dashboard generation.
- It is claimed to be the first benchmark for open-ended, task-grounded generation of such dashboards.
- The benchmark requires systems to generate both a dashboard and a replayable interaction trajectory.
- A browser executor replays the trajectory to produce visual and execution evidence.
- A VLM judge compares candidates using the evidence.
- Bradley–Terry aggregation produces the leaderboard.
- The judge is distilled into the open-weight DashJudge-8B model.
- The paper is available on arXiv with ID 2608.10567.
Entities
Institutions
- arXiv