r/LocalLLaMA
· Communities
10% faster decode with Q4_K MTP draft model with Gemma 4 31b
(Disclaimer: I am a noob and don’t know what I am doing) Gemma 4 31b unsloth/gemma-4-31B-it-qat-GGUF I took the f16 MTP draft model and quantised it to Q4_K (instead of Q4_0 of unsloth) and gained around 10% in decode: from 65TPs to 72TPs. Dual 3090, split mode layer. Draft KV to Q4_0 Anyone has the same experience or