Skip to content
arXiv cs.CV · Papers

Reasoning-Guided Part-Level Visual Grounding via Reinforcement Learning

arXiv:2607.15374v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) ground whole objects well from free-form language queries, but they struggle when the query names a part rather than the object. We trace this to a missing object-part hierarchy, since parts are localized in the same single step us