Skip to content
arXiv cs.CL · Papers

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs

arXiv:2607.28640v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) should generate consistent responses given semantically equivalent inputs across modalities. However, we observe a systematic discrepancy in model predictions under such cross-modal variations. Specifically, we define the modality