Skip to content
arXiv cs.CV · Papers

MMDiff: Extending Diffusion Transformers for Multi-Modal Generation

arXiv:2606.16673v2 Announce Type: replace Abstract: Diffusion transformers have demonstrated remarkable generative capabilities, yet the rich perceptual representations computed across their denoising trajectory are discarded once the content is rendered. We present MMDiff, a framework that transforms a frozen diffusio