arXiv cs.CV
· Papers
From Region Arrival to Instance-Level Grounding in Vision-and-Language Navigation
arXiv:2607.03792v1 Announce Type: cross Abstract: Vision-and-Language Navigation (VLN) agents may satisfy conventional success criteria while still failing to establish reliable object-level grounding, because current evaluation protocols mainly reward stopping within a 3-meter radius and largely ignore the agent's fin