arXiv cs.CL
· Papers
Keyless Attention: Value-Space Routing and Value-Only Caching for Efficient Transformers
arXiv:2606.21848v2 Announce Type: replace Abstract: Transformer architectures form the foundation of modern natural language processing, making it crucial to address the efficiency and scalability limitations of the standard QKV attention mechanism. The Key-Value (KV) cache is a major bottleneck during long-context inf