X · @teortaxesTex
· X / Twitter
> t's actually only equivalent at pass@3, sol is much more consistent Reminder that even GLM 5.2 and Kimi K3 are nowhere near RL-maxxed enough. They a…
> t's actually only equivalent at pass@3, sol is much more consistentReminder that even GLM 5.2 and Kimi K3 are nowhere near RL-maxxed enough. They are far from their capability ceiling. GPT, Grok, Opus are "more consistent" because of that. Just more steps neededpilvar (Philippe Dourassov): @daveaitel It's actually on