X Β· @teortaxesTex
Β· X / Twitter
RT Yu Zhang ππ: K3 has now crossed the 1M context-length barrier, and DeepSeek's sparse attn has done the same. But what architecture will take …
RT Yu Zhang ππK3 has now crossed the 1M context-length barrier, and DeepSeek's sparse attn has done the same. But what architecture will take us to 5M, 10M, or even longer? I'd always argue that fixed-state linear attn, especially GDN/KDA, is highly competitive here. Hybrid designs are scalable and extrapolate well to