Skip to content
arXiv cs.CV · Papers

Show Me Examples: Inferring Visual Concepts from Image Sets

arXiv:2607.02402v3 Announce Type: replace Abstract: Vision-language models (VLMs) can follow complex textual instructions, yet they struggle to reason from purely visual context. In particular, current models fail to infer shared concepts from sets of example images and apply them to new inputs. We introduce Visual Con