Skip to content
arXiv cs.CV · Papers

PixelEyes: Decoupling Perception and Reasoning for Pinpoint Visual Evidence Seeking

arXiv:2607.00115v1 Announce Type: new Abstract: This paper explores multi-turn visual reasoning and observes that MLLMs repeatedly fail to localize the target, leading to long, redundant trajectories. We attribute this failure to the entanglement of reasoning and perception within a single model, the MLLM reasons and l