Skip to content
arXiv cs.CV · Papers

Masked Visual Actions for Unified World Modeling

arXiv:2607.19343v1 Announce Type: new Abstract: Video models absorb rich priors over how the visual world moves, interacts, and responds to contact, making them promising substrates for robotic world modeling. The central challenge is how to communicate action to such models in a form aligned with the visual space in w