arXiv cs.LG
· Papers
rePIRL: Learn PRM with Inverse RL for LLM Reasoning
arXiv:2602.07832v3 Announce Type: replace Abstract: Process rewards have been widely used in deep reinforcement learning to improve training efficiency, reduce variance, and prevent reward hacking. In LLM reasoning, existing works also explore various solutions for learning effective process reward models (PRM) with or