r/LocalLLaMA
· Communities
llama.cpp MTP speculative simplified for July 2026 big wins on dense models, underwhelming on MoE
Wanted to consolidate where things actually stand now that the dust has settled on speculative decoding in llama.cpp, since the discourse a few months back was pretty scattered. The short version: native MTP (multi token prediction) support landed via --spec-type draft-mtp, letting models use their own built in MTP hea