ARTFEED — Contemporary Art Intelligence

FairGap Benchmark Exposes Hidden Fairness Gaps in LLM Recommenders

ai-technology · 2026-08-11

FairGap is a new benchmark that challenges the idea that consistent recommendations from LLM-based systems mean their internal processes are stable. Developed by researchers and detailed in arXiv:2608.08284, it evaluates recommendation fairness on two fronts: observable output shifts (OBS) and hidden representation shifts (IBS). By using controlled counterfactual identity probes looking at gender, age, and race, it introduces a concept called Representation-Output Alignment (ROA) to analyze these shifts. When applied to six open-weight LLM families in three different areas, FairGap reveals significant hidden-output discrepancies, with ROA rarely exceeding 0.22. Interestingly, many users see stable outputs while internal changes occur, a problem overlooked in output-only audits. The study also suggests that activation steering could lessen IBS by up to 8 times. You can find the full research on arXiv with the ID 2608.08284.

Key facts

  • FairGap is the first benchmark to jointly evaluate recommendation fairness at observable output shift (OBS) and hidden representation shift (IBS).
  • The benchmark uses controlled counterfactual identity probes across gender, age, and race.
  • Representation-Output Alignment (ROA) summarizes the relationship between OBS and IBS.
  • Quadrant diagnostics identify user-level hidden-output mismatch.
  • Applied to six open-weight LLM families across three domains.
  • ROA rarely exceeds 0.22, indicating pervasive hidden-output decoupling.
  • A non-negligible user population shows stable outputs despite substantial internal shifts.
  • Activation steering reduces IBS by up to 8x.
  • The paper is available on arXiv:2608.08284.

Entities

Institutions

  • arXiv

Sources