Skip to content
r/LocalLLaMA · Communities

Special Architecture in AFM3 20B: Instruction Following Pruning

https://openreview.net/forum?id=juARG7yu4P This is a model designed to activate ~20% of active MLP layers. It is also an MoE so it has some sparsity built-in. It's trained from scratch to use the same experts per prompt, not per token or switching per layer. Around two-thirds of a model's active parameters are FFN/MLP