Skip to content
arXiv cs.LG · Papers

Inference-Time Policy Alignment for Fair Reinforcement Learning

arXiv:2608.00175v1 Announce Type: new Abstract: Deep reinforcement learning (RL) agents achieve strong performance by optimizing scalar reward functions. However, once deployed, the policies of these RL agents are often rigid and costly to adapt to new performance criteria. For instance, an agent trained to maximize ex