X · @teortaxesTex
· X / Twitter
Interesting detail that NVIDIA's paper on LatentMoE already used a hypothetical Kimi-K2-LatentMoE as an example
Interesting detail that NVIDIA's paper on LatentMoE already used a hypothetical Kimi-K2-LatentMoE as an exampleSemiAnalysis: While it is true that Kimi K3 uses Kimi Delta Linear Attention (KDA) in 3 out of every 4 layers and that KDA reduces KV-cache transfer bandwidth by up to 10x compared with comparable full global-