arXiv cs.CV
· Papers
ETC: Extreme Token Compression via Task-aware Visual Information Distillation in VLMs
arXiv:2606.00543v2 Announce Type: replace Abstract: In Vision-Language Models (VLMs), high-resolution images produce a large number of visual tokens, resulting in high computational costs and KV-cache overhead during inference. To address this problem, we propose an Extreme Token Compression (ETC) framework that minimi