Skip to content
arXiv cs.CV · Papers

WorldPack: Dynamic Frame Compression for Long-context Video World Modeling

arXiv:2512.02473v2 Announce Type: replace Abstract: Video world models have attracted significant attention for their ability to produce high-fidelity future visual observations conditioned on past observations and navigation actions. However, achieving temporally and spatially consistent generation over long horizons