arXiv cs.CL
· Papers
Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model
arXiv:2607.24904v1 Announce Type: cross Abstract: Standard vision-language models (VLMs) suffer from Moravec's paradox: they excel at complex offline visual reasoning but struggle with simple streaming perception tasks and process them inefficiently. We present Mage-VL, an efficient codec-native streaming foundation mo