Skip to content
r/LocalLLaMA · Communities

Is there a point where models just cannot get any smaller without losing intelligence?

DeepSeek V4 Flash got me thinking... We keep seeing smaller models get way better. A model at a certain parameter count today can be much smarter than a model of the same size from a year or two ago. Better training, better data, better architectures, distillation, MoE, and all of that seem to let companies squeeze mor