Skip to content
X · @teortaxesTex · X / Twitter

Honestly this makes the whole benchmark look even more absurd. Grok 4.20 over Opus 4.8 (max), Kimi K2.5 > GLM 5.2 and Opus 4.7, Opus 4.6 down in the d…

Honestly this makes the whole benchmark look even more absurd. Grok 4.20 over Opus 4.8 (max), Kimi K2.5 > GLM 5.2 and Opus 4.7, Opus 4.6 down in the dumps below Grok 4… what is going on here? Sounds like it's super sensitive to lab priorities in this domain.prinz: Added to prinzbench: GLM-5.2.This is a slop model that