Inference-Time Policy Alignment for Fair Reinforcement Learning
A new arXiv paper (2608.00175v1) introduces a method for aligning pretrained reinforcement learning (RL) policies with fairness objectives at inference time, without updating the base policy's parameters. The approach, inspired by inference-time alignment in large language models, formalizes fairness alignment as a policy shaping problem and proposes a multiplicative policy shaping framework that adjusts action probabilities to meet welfare-based fairness criteria. This addresses the rigidity of RL policies that are optimized for scalar rewards and cannot accommodate previously unknown stakeholder preferences. Existing fairness approaches require complete retraining under a fairness-oriented metric, which is costly and assumes preferences are known a priori. The proposed method allows for steering a pretrained policy toward fairness objectives on the fly, making it adaptable to new performance criteria. The paper is available on arXiv and was announced as a cross-type submission.
Key facts
- Paper arXiv:2608.00175v1
- Announce type: cross
- Focus: inference-time policy alignment for fairness in RL
- Inspired by inference-time alignment in large language models
- Proposes multiplicative policy shaping framework
- Does not update base policy parameters
- Addresses rigidity of RL policies to new performance criteria
- Existing fairness approaches require complete retraining
Entities
Institutions
- arXiv