VarRate: Training-Free Variable-Rate KV Cache Compression for Long-Context LLMs
arXiv:2607.15498v1 Announce Type: new Abstract: The key-value (KV) cache is the main memory bottleneck in long-context large language model (LLM) inference. Two leading training-free families are…