X · @teortaxesTex
· X / Twitter
far as I can tell, K3 uses the latent bottleneck not to reduce routed traffic, but to spend the saved width on twice as many active experts. Its per-t…
far as I can tell, K3 uses the latent bottleneck not to reduce routed traffic, but to spend the saved width on twice as many active experts. Its per-token expert dispatch volume is exactly unchanged from K2—16 × 3,584 = 8 × 7,168.