New Framework Audits LLM Benchmarks at Sample Level
A recent preprint on arXiv (2607.28801) presents a meta-evaluation framework focused on datasets, which audits benchmark datasets for large language models (LLMs) at the sample level. This framework analyzes samples based on five latent dimensions: cognitive and knowledge demands, language and content quality, task properties, context, and ethics, safety, and fairness. The authors utilized this framework on five key benchmarks—MMLU, ARC, WinoGrande, HellaSwag, and TruthfulQA—uncovering notable internal variability that aggregate accuracy scores overlook. These annotations facilitate the criterion-driven organization of composite benchmark subsets across datasets, allowing for targeted assessments of specific model capabilities like reasoning depth or ethical sensitivity. The preprint, identified as arXiv:2607.28801v1, redefines benchmark evaluation by viewing benchmarks as diverse collections rather than singular tasks.
Key facts
- The framework audits benchmarks at the sample level.
- Five latent dimensions are used for annotation.
- The benchmarks annotated include MMLU, ARC, WinoGrande, HellaSwag, and TruthfulQA.
- The analysis reveals internal heterogeneity not captured by aggregate scores.
- Annotations enable orchestration of composite benchmark subsets.
- Targeted evaluation of capabilities like Reasoning Depth or Ethical Sensitivity is supported.
- The preprint is available on arXiv with ID 2607.28801.
- The announcement type is 'cross'.
Entities
Institutions
- arXiv