Skip to content
arXiv cs.CV · Papers

D3VL: Understanding Driving Scenes from 3D Time Series Data and Video with Language Models

arXiv:2607.19528v1 Announce Type: new Abstract: Recent advances in Multimodal Large Language Models (MLLMs) have triggered the development of end-to-end MLLMs for autonomous driving. However, the main emphasis to date has been for MLLMs using 2D images and videos. In contrast, this paper considers MLLM effectiveness us