Skip to content
arXiv cs.CV · Papers

Chain of Spatial Thoughts: Modality-Agnostic Spatial Grounding for Vision Language Models

arXiv:2608.10278v1 Announce Type: new Abstract: Spatial understanding is fundamental to embodied intelligence, underpinning applications such as robotic manipulation, embodied navigation, and autonomous driving. Although recent vision-language models (VLMs) have achieved impressive performance on spatial reasoning benc