X · @teortaxesTex
· X / Twitter
Interesting I guess we should say this is a decent amount of progress given the interval between 3.5 and 3.6 But, guys… you're going to need a bigger…
InterestingI guess we should say this is a decent amount of progress given the interval between 3.5 and 3.6But, guys… you're going to need a bigger modelLogan Kilpatrick: @MarcosHernanz we pushed 3.6 Flash on agentic use cases for real world tasks, AA mostly covers reasoning benchmarks which is why it didn’t change, we