Skip to content
X · @teortaxesTex · X / Twitter

> t's actually only equivalent at pass@3, sol is much more consistent Reminder that even GLM 5.2 and Kimi K3 are nowhere near RL-maxxed enough. They a…

> t's actually only equivalent at pass@3, sol is much more consistentReminder that even GLM 5.2 and Kimi K3 are nowhere near RL-maxxed enough. They are far from their capability ceiling. GPT, Grok, Opus are "more consistent" because of that. Just more steps neededpilvar (Philippe Dourassov): @daveaitel It's actually on