X · @teortaxesTex
· X / Twitter
> almost consistently outperforms GRPO is a strong, extensible baseline and I suspect that people dunking on it (and DeepSeek) are motivated by person…
> almost consistently outperformsGRPO is a strong, extensible baseline and I suspect that people dunking on it (and DeepSeek) are motivated by personal mathematical aesthetics and not so much sample efficiency. V4 vs 5.2 isn't a drama about wrong algorithmic betsGrad: New GLM paper on the PPO algo they use for GLM 5.2C