arXiv cs.CV
· Papers
PixelEyes: Decoupling Perception and Reasoning for Pinpoint Visual Evidence Seeking
arXiv:2607.00115v1 Announce Type: new Abstract: This paper explores multi-turn visual reasoning and observes that MLLMs repeatedly fail to localize the target, leading to long, redundant trajectories. We attribute this failure to the entanglement of reasoning and perception within a single model, the MLLM reasons and l