Skip to content
X · @teortaxesTex · X / Twitter

RT wh: Re The main change is to get rid of group-wise sampling in favor of single-rollout which means we now need a value model again to reduce varian…

RT whRe The main change is to get rid of group-wise sampling in favor of single-rollout which means we now need a value model again to reduce variance. Value model specifics:1) Value network takes 2 gradient steps per batch. This is shown to train a more accurate Value model mainly because the value model requires more