arXiv cs.CV
· Papers
MMDiff: Extending Diffusion Transformers for Multi-Modal Generation
arXiv:2606.16673v2 Announce Type: replace Abstract: Diffusion transformers have demonstrated remarkable generative capabilities, yet the rich perceptual representations computed across their denoising trajectory are discarded once the content is rendered. We present MMDiff, a framework that transforms a frozen diffusio