Skip to content
arXiv cs.LG · Papers

RGB-Pointmap Pretraining for Unified 3D Scene Understanding

arXiv:2604.02546v3 Announce Type: replace-cross Abstract: Pretraining 3D encoders through alignment with Contrastive Language-Image Pre-training (CLIP) has emerged as a promising direction for learning generalizable representations for 3D scene understanding. In this paper, we propose UniScene3D, a transformer-based fr