Skip to content
arXiv cs.CV · Papers

SportD: How do VLMs physically strategize?

arXiv:2607.14616v3 Announce Type: replace-cross Abstract: Vision-language models (VLMs) can describe a scene, but can they act well within one? We study whether VLMs can make sound strategic decisions, using soccer as an objective testbed with quantifiably-valued actions. We introduce SportD, a dataset and evaluation c