Skip to content
arXiv cs.CV · Papers

Parameter-Efficient CLIP Adaptation for 3D Understanding via Unified Tokenization

arXiv:2505.18819v2 Announce Type: replace Abstract: Vision-language models, such as CLIP, encode rich semantic knowledge through large-scale image-text pretraining. Reusing these models for 3D understanding is highly desirable, because 3D-text pairs and dense point-level annotations are far scarcer and more difficult t