Skip to content
X · @teortaxesTex · X / Twitter

RT gum: gave fable a shot at squeezing @AlpinDale's qwen3-0.6B megakernel further. It got 1,398 tok/s bf16 which is close to 5090s theoretical max. 32…

RT gumgave fable a shot at squeezing @AlpinDale's qwen3-0.6B megakernel further.It got 1,398 tok/s bf16 which is close to 5090s theoretical max. 32 of the 128 persistent blocks now do nothing but stream the next layers weights into L2.Alpin: New blogpost. I wanted to see how fast I can go with a 0.6B model on the RTX 5