arXiv cs.CV
· Papers
Reasoning-Guided Part-Level Visual Grounding via Reinforcement Learning
arXiv:2607.15374v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) ground whole objects well from free-form language queries, but they struggle when the query names a part rather than the object. We trace this to a missing object-part hierarchy, since parts are localized in the same single step us