We tested Deepseek v4 flash, GLM 5.2, and Kimi K3 on hard agentic tasks, and DeepSeek just crushed
https://preview.redd.it/uybjxyypj7hh1.png?width=1200&format=png&auto=webp&s=8293e8da332a14920b335caee52473762c09530d We just tested the big 3 of open-weight models on our internal evals. The tasks include long-running workflows involving multiple applications (Pagerduty, Gmail, HubSpot, Airtable,…