arXiv cs.LG
· Papers
Safe Inference-Time Alignment via Lagrangian Reward Augmentation
arXiv:2607.02781v1 Announce Type: new Abstract: Inference-time alignment steers a frozen language model during decoding using auxiliary reward signals, avoiding the cost of repeated weight updates. However, existing inference-time alignment methods typically optimize a single scalar score, so explicit safety constraint