Skip to content
X · @togethercompute · X / Twitter

We analyzed Kimi K3 and GPT-5.6 Sol on DeepSWE. A Kimi-first cascade with test-suite verification outperformed Sol alone at a lower cost per completed…

We analyzed Kimi K3 and GPT-5.6 Sol on DeepSWE. A Kimi-first cascade with test-suite verification outperformed Sol alone at a lower cost per completed task.