Skip to content
r/LocalLLaMA · Communities

microsoft/Mage-VL · Hugging Face – An Efficient Codec-Native Streaming Multimodal Foundation Model

Mage-VL is a codec-native, proactive-streaming multimodal foundation model for image and video understanding, whose visual encoder is trained entirely from scratch at a compact 4B scale. It targets a modern Moravec's paradox of VLMs — strong at complex offline reasoning, yet slow and compute-heavy on simple real-time s