HF Daily Papers
· Papers
SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them
Vision-language models (VLMs) are increasingly used in embodied agents to interpret visual inputs, reason about spatial relationships, and make task-level decisions based on that reasoning. However, a fundamental capability mismatch remains: general VLMs can reason about the overall task but often m