Skip to content
arXiv cs.LG · Papers

Learning to Reason Efficiently with Discounted Reinforcement Learning

arXiv:2510.23486v3 Announce Type: replace Abstract: Large reasoning models (LRMs) often consume excessive tokens, inflating computational cost and latency. More broadly, in goal reaching sequential decision problems we often want to reach the goal quickly, and LRM reasoning can be viewed through this lens. We challenge