RLHF Procedural Fairness Failures from Preference Averaging
A recent study posted on arXiv under the identifier 2608.10126 highlights shortcomings in the procedural fairness of Reinforcement Learning from Human Feedback (RLHF). The researchers point out that the conventional approach amalgamates varied preferences into a single reward model, resulting in the overshadowing of minority opinions by majority preferences. They emphasize that procedural fairness involves maintaining clear preference signals in reward modeling, a goal not met by standard RLHF methods. To tackle this issue, the authors propose a new framework called Preference-Aware RLHF (PA-RLHF), which enhances alignment accuracy from 46.9% to 67.9% and significantly narrows the fairness gap among aligned groups.
Key facts
- Paper arXiv:2608.10126
- RLHF aggregates heterogeneous preferences into a single reward model
- Preference averaging causes procedural fairness failures
- Majority preference groups dominate reward learning
- Minority preferences are systematically under-represented
- PA-RLHF separates optimization across preference modes
- PA-RLHF improves alignment accuracy from 46.9% to 67.9%
- Fairness gap reduced from 15.9 to 9.6 percentage points
Entities
—