Skip to content
r/LocalLLaMA · Communities

I gave DeepSeek V4 Flash basic vision by training a 40M connector on 100K examples

I wanted to find out whether a huge text-only MoE could be given basic vision without retraining the language model itself. The short answer is yes. I froze DeepSeek V4 Flash and a 417M-parameter MoonViT image encoder, then trained a 40.1M-parameter connector between them on 100,000 image-text examples. The completed N