X · @teortaxesTex
· X / Twitter
RT Xiuyu Li: RL with verifiable rewards propels reasoning past imitation, but it spreads the reward signal across every token in a trajectory, most of…
RT Xiuyu LiRL with verifiable rewards propels reasoning past imitation, but it spreads the reward signal across every token in a trajectory, most of which are routine and carry no decision.They select tokens adaptively using a relative surprisal index that flags the ones where the policy's choice genuinely diverged and