arXiv cs.CL
· Papers
Thought-Aware KV Cache Compaction for Reasoning via Adaptive Attention Matching
arXiv:2608.12331v1 Announce Type: new Abstract: Reasoning language models generate lengthy chain-of-thought (CoT) sequences whose key-value (KV) cache grows linearly and becomes a memory bottleneck during decoding. Existing compaction methods treat reasoning trajectories as flat token sequences and apply uniform compre