X · @teortaxesTex
· X / Twitter
Zephyr is probably right that they're saving resources for training, except I'm not sure about 3-4T. We see how absurdly strong Flash is. (or new dots…
Zephyr is probably right that they're saving resources for training, except I'm not sure about 3-4T. We see how absurdly strong Flash is. (or new dots). They might first train a V4-Clean, ablated simplified architecture; then scale it up in Q1 2027.V4s will be RL teachers anywayZephyr: IMO, they have finally stabilized