Skip to content
arXiv cs.CV · Papers

Pixel-Space Diffusion Transformers

arXiv:2607.17585v2 Announce Type: replace Abstract: Latent diffusion models (LDMs) enable efficient high-resolution image synthesis by denoising in a VAE-compressed latent space. However, fixed visual tokenizers can discard fine textures and structural details, while separate representation and diffusion training creat