Skip to content
arXiv cs.CV · Papers

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models

arXiv:2502.09696v3 Announce Type: replace Abstract: Large Multimodal Models (LMMs) exhibit shortfalls when interpreting images and, by some measures, have poorer spatial cognition than young children or animals. Despite this, they attain high scores on many popular visual benchmarks, with headroom rapidly eroded by mod