X · @teortaxesTex
· X / Twitter
RT Elliot Arledge: For those wondering why I use a Kimi Linear megakernel instead of Qwen 3.6, first look at the parameter counts. One is 35 billion, …
RT Elliot ArledgeFor those wondering why I use a Kimi Linear megakernel instead of Qwen 3.6, first look at the parameter counts. One is 35 billion, one is 48 billion, and they're both 3 billion active experts. So they're going to use the same amount of weights in total for, or roughly the same amount of weights for pre