New Benchmark MAS-HQ Evaluates AI Factuality vs. Compute Cost
A recent paper published on arXiv presents MAS-HQ (Multi-Agent System Hallucination Quest), a protocol designed to evaluate resource-efficient benchmarks for hallucination in large language models. The researchers contend that traditional static leaderboards, which assess factuality in isolation and disregard computational expenses, fail to differentiate truly superior systems from those that merely utilize more computational power. They illustrate a ranking shift: a brute-force Best-of-4 agent secures a higher raw factuality score (H-Score 0.9169 compared to 0.9103) but falls short in the cost-adjusted Q-Score (0.5169 versus 0.5217) while consuming approximately four times the tokens and latency. The protocol integrates any factuality detector, highlighting the balance between accuracy and computational costs.
Key facts
- Paper introduces MAS-HQ (Multi-Agent System Hallucination Quest) protocol
- Static leaderboards treat compute as free, masking cost-performance trade-offs
- Best-of-4 agent has higher H-Score (0.9169) but lower Q-Score (0.5169) than a more efficient system
- The efficient system has H-Score 0.9103 and Q-Score 0.5217
- Cost weight sensitivity is swept in the analysis
- Protocol wraps any factuality detector
- Published on arXiv under ID 2607.24063
- Research focuses on frontier models near top of factuality scale
Entities
Institutions
- arXiv