Skip to content
arXiv cs.LG · Papers

Annealed Softmax Greedy in Many-Armed Bayesian Bandits

arXiv:2605.31034v2 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards (RLVR) and group-based policy optimization methods such as GRPO update a stochastic policy by sampling multiple completions per prompt and increasing the policy's probability on those with higher reward, regularized by a