Skip to content
X · @teortaxesTex · X / Twitter

I'm confused by TB 3.0 looks like it measures general intelligence X general "agenticness", so both very strong and very harnessmaxxed models get ahea…

I'm confused by TB 3.0looks like it measures general intelligence X general "agenticness", so both very strong and very harnessmaxxed models get ahead. I don't think Opus 5 is really superior to Sol思维怪怪: Terminal-Bench 3.0 最新榜单换王。这个评测让 AI Agent 在真实终端环境里完成任务,最后直接检查结果是否正确。当前前三名分别是:1. Claude Opus 5 Max + mini-SWE-agent:43