arXiv cs.LG
· Papers
How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures
arXiv:2608.13267v1 Announce Type: cross Abstract: Existing vision-language model (VLM) benchmarks emphasize perception and reasoning accuracy (how well VLMs describe and reason about what they see in an image), with limited attention to behavioral reliability under uncertainty (how they behave when visual evidence is m