Skip to content
r/LocalLLaMA · Communities

A 460M VLM gets first-token latency down to 0.3s on an iPhone by using only 64 visual tokens

A new 460M vision model called VisionPsy-Nano-460M-Flash is taking a slightly different approach to on-device VLM speed. Instead of making the language model much smaller, it reduces how much visual information reaches it. For a 512x512 image: - VisionPsy Flash: 64 visual tokens - LFM2.5-VL-450M: 256 - Qwen3.5-0.8B: 25