Skip to content
X · @teortaxesTex · X / Twitter

Damn, that's right. So: "≈48B" (thanks to N-Gram embedding, variable) active, 35T tokens. V4-tier, ≈8e24? Would be the biggest Chinese pretraining o…

Damn, that's right.So: "≈48B" (thanks to N-Gram embedding, variable) active, 35T tokens. V4-tier, ≈8e24? Would be the biggest Chinese pretraining on domestic hardware. Some strange "Superpods":> "our accelerators"> "up to 48 machines each"> below 80 GB HBM per accelerator > "The device offers limited HBM bandwidth but