arXiv cs.LG
· Papers
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning
arXiv:2511.02130v2 Announce Type: replace-cross Abstract: We propose Re-FORC, an adaptive reward prediction method that, given a query, enables prediction of the expected future rewards as a function of the number of future thinking tokens. Re-FORC trains a lightweight adapter on reasoning models, demonstrating improve