FairGap Benchmark Exposes Hidden Fairness Gaps in LLM Recommenders
FairGap is a new benchmark that challenges the idea that consistent recommendations from LLM-based systems mean their internal processes are stable. Developed by researchers and detailed in arXiv:2608.08284, it evaluates recommendation fairness on two fronts: observable output shifts (OBS) and hidden representation shifts (IBS). By using controlled counterfactual identity probes looking at gender, age, and race, it introduces a concept called Representation-Output Alignment (ROA) to analyze these shifts. When applied to six open-weight LLM families in three different areas, FairGap reveals significant hidden-output discrepancies, with ROA rarely exceeding 0.22. Interestingly, many users see stable outputs while internal changes occur, a problem overlooked in output-only audits. The study also suggests that activation steering could lessen IBS by up to 8 times. You can find the full research on arXiv with the ID 2608.08284.
Key facts
- FairGap is the first benchmark to jointly evaluate recommendation fairness at observable output shift (OBS) and hidden representation shift (IBS).
- The benchmark uses controlled counterfactual identity probes across gender, age, and race.
- Representation-Output Alignment (ROA) summarizes the relationship between OBS and IBS.
- Quadrant diagnostics identify user-level hidden-output mismatch.
- Applied to six open-weight LLM families across three domains.
- ROA rarely exceeds 0.22, indicating pervasive hidden-output decoupling.
- A non-negligible user population shows stable outputs despite substantial internal shifts.
- Activation steering reduces IBS by up to 8x.
- The paper is available on arXiv:2608.08284.
Entities
Institutions
- arXiv