r/LocalLLaMA
· Communities
MTP on MoE matters
Hey everyone, Back when MTP came available on llama.cpp, it seemed like the common consensus was that MTP didn't matter much for MoE models. After spending an evening running tests, I got some really decent performance increases out of Gemma4-26B-A4B-IT-QAT. From 88 t/s TG to 132 t/s TG Seems for me that n-max 3 min-p