ARTFEED — Contemporary Art Intelligence

SA-PPG: New Metric for Benchmark Contamination Mitigation

ai-technology · 2026-08-10

A recent study published on arXiv (2608.07341) introduces SA-PPG (Stratified Aggregate of Per-question Probability Gaps) as a refined metric for assessing contamination reduction in large language models. The researchers contend that the existing metric, G-AP (Gap of Aggregate Performance), has significant shortcomings. They highlight that binary correct/incorrect outputs fail to adequately represent performance on a per-question basis, and averaging prior to differencing can lead to the cancellation of both over- and under-suppression. Additionally, they note that uniform weighting per question can lead to tactics that skew solve probabilities toward the clean model's high-frequency values. SA-PPG calculates the solve probability for each question through sampling, compares it to the clean model, and aggregates results within groups. This paper is classified as a cross-type announcement and can be accessed at https://arxiv.org/abs/2608.07341.

Key facts

  • Paper arXiv:2608.07341 proposes SA-PPG metric
  • SA-PPG stands for Stratified Aggregate of Per-question Probability Gaps
  • G-AP (Gap of Aggregate Performance) is criticized as flawed
  • G-AP fails to characterize per-question performance
  • Averaging before differencing causes over- and under-suppression to cancel out
  • Uniform per-question weighting invites strategies to push solve probabilities
  • SA-PPG estimates solve probability per question by sampling
  • SA-PPG differences against clean model per question and aggregates within groups

Entities

Institutions

  • arXiv

Sources