X · @teortaxesTex
· X / Twitter
However, 27B dense Qwen is likely pretrained with much higher MFU than DSV4-Flash 13AB has been, so they are about equally costly in GPU-hours per 1T …
However, 27B dense Qwen is likely pretrained with much higher MFU than DSV4-Flash 13AB has been, so they are about equally costly in GPU-hours per 1T tokens (training)I wish we started seeing these figures again. V2, V3, gpt-oss, Nemotrons. Do you know any more relevant anchors?Tiezhen WANG: On that:27B dense is actual