r/LocalLLaMA
· Communities
Paper claims RL for reasoning only changes 1-3% of tokens, and they replicate the gains without RL at ~1000x less compute
submitted by /u/juanviera23 [link] [comments]
submitted by /u/juanviera23 [link] [comments]