XL-DocBench: New Benchmark for Extra-Long Document Understanding
Researchers have introduced XL-DocBench, a fully human-verified benchmark designed to evaluate large language models (LLMs) on extra-long document understanding tasks. The benchmark addresses the gap in existing evaluation methods, which typically focus on short-context or single-page question answering, by testing models on documents spanning up to 2,303 pages. XL-DocBench includes 1,519 retained questions across six professional domains, including compliance, clinical, financial, and engineering fields. Notably, 72.6% of the questions (1,103 examples) require evidence from multiple pages, and 36.6% (556 questions) involve tables, charts, or figures. The benchmark emphasizes traceability to specific evidence pages, reflecting the high cost of unsupported answers in professional workflows. This development is significant for advancing LLM capabilities in real-world document tasks, such as analyzing annual reports, regulations, clinical guidelines, and technical manuals. The benchmark was announced on arXiv with the identifier 2608.00036.
Key facts
- XL-DocBench is a fully human-verified benchmark for extra-long document understanding.
- It includes 1,519 retained questions from six professional domains.
- Contexts in the benchmark span up to 2,303 pages.
- 72.6% of questions (1,103 examples) use multiple evidence pages.
- 36.6% of questions (556) involve tables, charts, or figures.
- The benchmark focuses on traceability to specific evidence pages.
- It addresses gaps in existing benchmarks that focus on short-context or single-page QA.
- The benchmark was announced on arXiv with identifier 2608.00036.
Entities
—