X · @teortaxesTex
· X / Twitter
Damn, that's right. So: "≈48B" (thanks to N-Gram embedding, variable) active, 35T tokens. V4-tier, ≈8e24? Would be the biggest Chinese pretraining o…
Damn, that's right.So: "≈48B" (thanks to N-Gram embedding, variable) active, 35T tokens. V4-tier, ≈8e24? Would be the biggest Chinese pretraining on domestic hardware. Some strange "Superpods":> "our accelerators"> "up to 48 machines each"> below 80 GB HBM per accelerator > "The device offers limited HBM bandwidth but