arXiv cs.CV
· Papers
SceneActBench: Can Agents Act on the 3D Scenes They See?
arXiv:2607.22393v1 Announce Type: cross Abstract: Vision-language model (VLM) agents increasingly use tools to act on 3D scenes rather than only describe them. Existing 3D benchmarks score textual responses or single-object operations, leaving agent action on complete multi-object 3D scenes under evaluated. We present