Skip to content
X · @teortaxesTex · X / Twitter

RT Xiuyu Li: RL with verifiable rewards propels reasoning past imitation, but it spreads the reward signal across every token in a trajectory, most of…

RT Xiuyu LiRL with verifiable rewards propels reasoning past imitation, but it spreads the reward signal across every token in a trajectory, most of which are routine and carry no decision.They select tokens adaptively using a relative surprisal index that flags the ones where the policy's choice genuinely diverged and