Skip to content
arXiv cs.LG · Papers

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing

arXiv:2607.08497v1 Announce Type: cross Abstract: Recent unified multimodal models show a single architecture can jointly perform vision/language understanding and image generation/editing. However, they repeatedly feed all historical visual and textual inputs into a shared context window, limiting long-horizon multimo