Skip to content
arXiv cs.CV · Papers

StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling

arXiv:2507.05240v2 Announce Type: replace-cross Abstract: Vision-and-Language Navigation (VLN) in real-world settings requires agents to process continuous visual streams and generate actions with low latency grounded in language instructions. While Video-based Large Language Models (Video-LLMs) have driven recent prog