arXiv cs.CV
· Papers
Neural Introspection Gating for Adaptive KV-Cache Reuse in Vision-Language-Action Models
arXiv:2608.10824v1 Announce Type: cross Abstract: Vision-Language-Action(VLA) models map camera images and language instructions directly to motor commands through a single autoregressive transformer. In real-time control, they still spend substantial compute recomputing key-value(KV) representations for visual tokens