Skip to content
arXiv cs.CV · Papers

Visual Token Compression Enhances Robustness of MLLMs

arXiv:2607.22716v1 Announce Type: new Abstract: In this paper, we show for the first time that visual token pruning enhances the robustness of Multimodal Large Language Models (MLLMs), mitigating vulnerabilities such as jailbreak attacks and hallucinations. Given that vision and language modalities cannot be perfectly