Skip to content
arXiv cs.AI · Papers

Progressive Multimodal Alignment for Continual Instruction Tuning

arXiv:2607.26947v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) rely on a projector to align visual representations with the language embedding space, making it central to cross-modal understanding. In Multimodal Continual Instruction Tuning (MCIT), however, shifting visual distributi