arXiv cs.CV
· Papers
Dynamic Execution Commitment of Vision-Language-Action Models
arXiv:2605.11567v4 Announce Type: replace Abstract: Vision-Language-Action (VLA) models predominantly adopt action chunking, i.e., predicting and committing to a short horizon of consecutive low-level actions in a single forward pass, to amortize the inference cost of large-scale backbones and reduce per-step latency.