r/LocalLLaMA
· Communities
Why are MoE models so belittled?
E.g "Qwen 3.5 122B is just 10B active, so it's no where close to the dense 27B model" That is the main sentiment around here and it puzzles me. If a 122B is just worth 10B, then why does model providers bother creating an MoE model when they could've just released a dense 10B model? It sure is not that simple. I mean y