Skip to content
arXiv cs.CL · Papers

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model

arXiv:2607.24904v1 Announce Type: cross Abstract: Standard vision-language models (VLMs) suffer from Moravec's paradox: they excel at complex offline visual reasoning but struggle with simple streaming perception tasks and process them inefficiently. We present Mage-VL, an efficient codec-native streaming foundation mo