arXiv cs.CL
· Papers
TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs
arXiv:2607.28640v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) should generate consistent responses given semantically equivalent inputs across modalities. However, we observe a systematic discrepancy in model predictions under such cross-modal variations. Specifically, we define the modality