X · @emollick
· X / Twitter
You really need to benchmark models for your use case. As soon as judgements & decisions stack on top of each other, the differences between models am…
You really need to benchmark models for your use case.As soon as judgements & decisions stack on top of each other, the differences between models amplifies, and no standard benchmark will tell you that Gemini 3.1 is less worried about financial losses at a cafe than GPT-5.5Andon Labs: Gemini 3.1 Pro lost $6k running A