arXiv cs.CV
· Papers
Masked Visual Actions for Unified World Modeling
arXiv:2607.19343v1 Announce Type: new Abstract: Video models absorb rich priors over how the visual world moves, interacts, and responds to contact, making them promising substrates for robotic world modeling. The central challenge is how to communicate action to such models in a form aligned with the visual space in w