Skip to content
arXiv cs.CV · Papers

Jolia: Concept-Level Vision-Language Alignment for 3D CT Contrastive Learning

arXiv:2606.24570v2 Announce Type: replace Abstract: Vision-language contrastive pretraining has become the dominant recipe for 3D medical foundation models, leveraging the large volumes of paired scans and reports produced in clinical practice. However, medical images usually span dozens of organs, and radiological rep