HF Daily Papers
· Papers
Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models
Multimodal Large Language Models (MLLMs) achieve strong performance by integrating visual inputs with the rich priors of pretrained language models. However, they often fail on vision-centric tasks, especially when visual evidence conflicts with pretrained knowledge. We explore these failures separa