Label-Free Strategies Fail to Improve Accuracy in Multiple-Choice Benchmarks
A recent study published on arXiv (2608.11947) explores whether restricting large language models from accessing option labels during multiple-choice questions can eliminate positional bias and enhance accuracy. The research assesses two strategies that do not use labels: generation-then-matching and evaluating options independently. Unfortunately, neither method consistently boosts accuracy. An analysis indicates that the main issue lies in withholding options rather than in the matching process. The only setup that reliably aligns with the baseline involves presenting all options alongside an LLM matcher. Nonetheless, completely removing positional effects does not consistently lead to improvements in accuracy. These results challenge the efficacy of label-free methods in addressing order sensitivity in MCQ evaluations.
Key facts
- Paper arXiv:2608.11947 examines label-free strategies for multiple-choice benchmarks.
- Two strategies tested: generation-then-matching and scoring options in isolation.
- Neither strategy reliably improves accuracy.
- Withholding options is the bottleneck, not the matching step.
- Only configuration matching baseline: all options with LLM matcher.
- Eliminating positional influence does not reliably yield accuracy gains.
- MCQ scores conflate knowledge with sensitivity to option order.
- Study questions reliability of MCQ benchmarks for model knowledge.
Entities
Institutions
- arXiv