Skip to content
X · @teortaxesTex · X / Twitter

> almost consistently outperforms GRPO is a strong, extensible baseline and I suspect that people dunking on it (and DeepSeek) are motivated by person…

> almost consistently outperformsGRPO is a strong, extensible baseline and I suspect that people dunking on it (and DeepSeek) are motivated by personal mathematical aesthetics and not so much sample efficiency. V4 vs 5.2 isn't a drama about wrong algorithmic betsGrad: New GLM paper on the PPO algo they use for GLM 5.2C