ARTFEED — Contemporary Art Intelligence

Inference-Time Policy Alignment for Fair Reinforcement Learning

ai-technology · 2026-08-04

A new arXiv paper (2608.00175v1) introduces a method for aligning pretrained reinforcement learning (RL) policies with fairness objectives at inference time, without updating the base policy's parameters. The approach, inspired by inference-time alignment in large language models, formalizes fairness alignment as a policy shaping problem and proposes a multiplicative policy shaping framework that adjusts action probabilities to meet welfare-based fairness criteria. This addresses the rigidity of RL policies that are optimized for scalar rewards and cannot accommodate previously unknown stakeholder preferences. Existing fairness approaches require complete retraining under a fairness-oriented metric, which is costly and assumes preferences are known a priori. The proposed method allows for steering a pretrained policy toward fairness objectives on the fly, making it adaptable to new performance criteria. The paper is available on arXiv and was announced as a cross-type submission.

Key facts

  • Paper arXiv:2608.00175v1
  • Announce type: cross
  • Focus: inference-time policy alignment for fairness in RL
  • Inspired by inference-time alignment in large language models
  • Proposes multiplicative policy shaping framework
  • Does not update base policy parameters
  • Addresses rigidity of RL policies to new performance criteria
  • Existing fairness approaches require complete retraining

Entities

Institutions

  • arXiv

Sources