Skip to content
arXiv cs.LG · Papers

Output-Aware Rotation for INT2 KV-Cache Quantization

arXiv:2608.02691v1 Announce Type: new Abstract: The key-value (KV) cache has become a major memory and bandwidth bottleneck in long-context large language model inference, making ultra-low-bit quantization increasingly important. However, existing rotation-based INT2 methods optimize cache statistics or proxy errors be