Skip to content
r/LocalLLaMA · Communities

Paper claims RL for reasoning only changes 1-3% of tokens, and they replicate the gains without RL at ~1000x less compute

submitted by /u/juanviera23 [link] [comments]