Skip to content
arXiv cs.CV · Papers

From Alignment to Synthesis Contrastive Volumetric Grounding for Text-to-CT Generation

arXiv:2506.00633v3 Announce Type: replace Abstract: Generating semantically controllable 3D CT volumes from radiology reports requires more than a rich text encoder, it requires vision-language alignment grounded in volumetric space. Existing Text-to-CT approaches condition generation on encoders pretrained with langua