r/LocalLLaMA
· Communities
Speculative decoding with deepseek v4 flash 0731?
Has anyone figured out how to enable speculative decoding with deepseek v4 flash 0731 on llamacpp? I’m on the right release for llamacpp (b10228 or earlier) and running am17an’s draft model with unsloth’s UD-Q8 model. Running into a lot of issues however. My llamacpp command: CUDA_DEVICE_ORDER=PCI_BUS_ID ~/llama.cpp/