Skip to content
arXiv cs.CV · Papers

Depth-Wise Probing and Pruning of the Planning Token in a Driving Vision-Language-Action Model

arXiv:2608.07361v1 Announce Type: cross Abstract: Vision-language-action (VLA) models route driving decisions through a deep language model, but it is unclear how much of that depth the action itself requires. We study a representative driving VLA whose entire plan is carried by a single planning token that a generativ