Skip to content
arXiv cs.CV · Papers

Which Modality Decides? Counterfactual Modality Attribution for Multimodal LLMs

arXiv:2608.00076v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) increasingly support high-stakes decision making by combining complementary information from images and text. While existing explainability methods identify influential image regions or text tokens, they cannot answer a fundamental