Skip to content
arXiv cs.AI · Papers

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin

arXiv:2608.06411v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) achieve strong performance across diverse vision-language tasks, but their efficiency is limited by the cost of processing numerous visual tokens. Visual token pruning can reduce this cost, but requires accurate token importance es