arXiv cs.LG
· Papers
Learning to Reason Efficiently with Discounted Reinforcement Learning
arXiv:2510.23486v3 Announce Type: replace Abstract: Large reasoning models (LRMs) often consume excessive tokens, inflating computational cost and latency. More broadly, in goal reaching sequential decision problems we often want to reach the goal quickly, and LRM reasoning can be viewed through this lens. We challenge