arXiv cs.CV
· Papers
Can Vision-Language Models Assess Proxemic Risk from Egocentric Robot Images?
arXiv:2608.12515v1 Announce Type: new Abstract: Assessing proxemic danger from a robot's egocentric perspective is critical for safe embodied navigation in human environments and requires both visual and contextual reasoning. We evaluate three opensource vision-language models (VLMs) (textit{InternVL}, textit{Qwen-VL