X · @teortaxesTex
· X / Twitter
RT Igor Kotenkov: It's interesting how GPT-5.5 behaves like a 🔨mere tool🔨, just doing the work to satisfy the tests, while Anthropic models win …
RT Igor KotenkovIt's interesting how GPT-5.5 behaves like a 🔨mere tool🔨, just doing the work to satisfy the tests, while Anthropic models win if scoring includes "taste"/bloatness of the code/etc.(also note GLM scores 🫥)Braden Hancock: New gold standard benchmark for measuring agentic coding abilities just dropped: Sen