arXiv cs.CV
· Papers
D3VL: Understanding Driving Scenes from 3D Time Series Data and Video with Language Models
arXiv:2607.19528v1 Announce Type: new Abstract: Recent advances in Multimodal Large Language Models (MLLMs) have triggered the development of end-to-end MLLMs for autonomous driving. However, the main emphasis to date has been for MLLMs using 2D images and videos. In contrast, this paper considers MLLM effectiveness us