arXiv cs.CV
· Papers
MaskAttn-SDXL: Controllable Region-Level Text-To-Image Generation
arXiv:2509.15357v3 Announce Type: replace Abstract: Diffusion models have achieved strong results in text-to-image generation, but important limitations remain as prompts become more structured and multi-object. On the architecture side, U-Net backbones are efficient and stable, yet their locality makes global coordina