LLM Panel Aggregation: Pivotal Votes Only Where Verification Helps
A recent study published on arXiv (2608.06940) examines the effectiveness of panels of LLM judges, focusing on the problem of correlated errors among them. The findings indicate that while a group of nine judges yields information comparable to that of two independent judges, aggregation only marginally reduces the error gap. An alternative approach—utilizing a signal from a different source of evidence, like running a test suite—demonstrated no significant change in the effective-vote count at scale (-0.04, 95% CI [-0.10, +0.02]). The authors contend that aggregate dependence and conditional decision utility are separate issues. They highlight that only decisions with a one-vote margin can be influenced by single-ballot substitution, with panel error rates increasing at these critical points, achieving accuracy improvements of +10.4 to +23.3 percentage points in three main configurations, and no improvements elsewhere. The paper was released on arXiv under ID 2608.06940v1.
Key facts
- Paper ID: arXiv:2608.06940v1
- LLM judge panels exhibit highly correlated errors
- Nine judges provide roughly the effective information of two independent ones
- Aggregation closes only a small fraction of the gap
- Adding a test suite signal produced no distinguishable change in effective-vote count (-0.04, 95% CI [-0.10, +0.02])
- Only decisions with a one-vote margin can change with single-ballot substitution
- Accuracy gains concentrate on pivotal queries: +10.4 to +23.3 percentage points across three headline configurations
- Accuracy gain is exactly zero elsewhere
Entities
Institutions
- arXiv