Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model
arXiv:2607.24904v1 Announce Type: cross Abstract: Standard vision-language models (VLMs) suffer from Moravec's paradox: they excel at complex offline visual reasoning but struggle with simple streaming perception…