arXiv cs.LG
· Papers
Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms
arXiv:2608.04074v1 Announce Type: new Abstract: Long-context LLM decoding reads the key-value (KV) cache at every step. Loading it takes longer than computing attention over it, so throughput is bandwidth-bound. Hence, reducing the cache size can raise both decoding speed and serving capacity. The challenge is to reduc