X · @teortaxesTex
· X / Twitter
Sounds about right. Not even shocking. We know of open models with similar compute footprint and >40T pretraining tokens (gpt-oss, MiMo). This approac…
Sounds about right. Not even shocking. We know of open models with similar compute footprint and >40T pretraining tokens (gpt-oss, MiMo). This approach is also why Demis says they don't have compute for open weights.Kek. Gemma 4 project has comparable footprint to GLM 5.2/DSV4.elie: seems like gemma 4 was trained on MU