ARTFEED — Contemporary Art Intelligence

LivingArena: Peer-Probing Framework for LLM Evaluation

ai-technology · 2026-07-29

A new automated evaluation framework called LivingArena uses peer-probing to assess frontier LLMs, addressing contamination and saturation issues in static benchmarks. In this framework, models take turns proposing questions to exploit opponents' knowledge boundaries, with a judge panel validating objective answers. The approach yields contamination-resistant evaluations of ten frontier LLMs.

Key facts

  • LivingArena is an automated, contamination-resistant evaluation framework for LLMs.
  • Models propose questions to opponents, aiming to identify knowledge gaps.
  • A judge panel of strong models validates question objectivity.
  • Questioners are rewarded when opponents fail; answerers are rewarded otherwise.
  • The framework evaluates ten frontier LLMs.
  • Static benchmarks suffer from contamination and saturation.
  • Human preference is subjective for evaluation.
  • The paper is available on arXiv with ID 2607.24780.

Entities

Institutions

  • arXiv

Sources