Skip to content
arXiv cs.CL · Papers

Keyless Attention: Value-Space Routing and Value-Only Caching for Efficient Transformers

arXiv:2606.21848v2 Announce Type: replace Abstract: Transformer architectures form the foundation of modern natural language processing, making it crucial to address the efficiency and scalability limitations of the standard QKV attention mechanism. The Key-Value (KV) cache is a major bottleneck during long-context inf