SA-PPG: New Metric for Benchmark Contamination Mitigation
A recent study published on arXiv (2608.07341) introduces SA-PPG (Stratified Aggregate of Per-question Probability Gaps) as a refined metric for assessing contamination reduction in large language models. The researchers contend that the existing metric, G-AP (Gap of Aggregate Performance), has significant shortcomings. They highlight that binary correct/incorrect outputs fail to adequately represent performance on a per-question basis, and averaging prior to differencing can lead to the cancellation of both over- and under-suppression. Additionally, they note that uniform weighting per question can lead to tactics that skew solve probabilities toward the clean model's high-frequency values. SA-PPG calculates the solve probability for each question through sampling, compares it to the clean model, and aggregates results within groups. This paper is classified as a cross-type announcement and can be accessed at https://arxiv.org/abs/2608.07341.
Key facts
- Paper arXiv:2608.07341 proposes SA-PPG metric
- SA-PPG stands for Stratified Aggregate of Per-question Probability Gaps
- G-AP (Gap of Aggregate Performance) is criticized as flawed
- G-AP fails to characterize per-question performance
- Averaging before differencing causes over- and under-suppression to cancel out
- Uniform per-question weighting invites strategies to push solve probabilities
- SA-PPG estimates solve probability per question by sampling
- SA-PPG differences against clean model per question and aggregates within groups
Entities
Institutions
- arXiv