Skip to content
arXiv cs.LG · Papers

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning

arXiv:2511.02130v2 Announce Type: replace-cross Abstract: We propose Re-FORC, an adaptive reward prediction method that, given a query, enables prediction of the expected future rewards as a function of the number of future thinking tokens. Re-FORC trains a lightweight adapter on reasoning models, demonstrating improve