Skip to content
arXiv cs.AI · Papers

See like a Robot: Robot-Centric Pointmaps for Vision-Language-Action Models

arXiv:2607.11498v1 Announce Type: cross Abstract: Vision-language-action (VLA) models predict robot actions from visual observations and language instructions. These actions are defined in the robot's own 3D coordinate frame, yet most VLAs observe the scene in the camera frame, creating a frame mismatch between where t