r/LocalLLaMA
· Communities
I benchmarked Unsloth’s Qwen3.6-27B NVFP4 on 1x/2x 5090s. MTP is great until it really isn’t.
I've been using the GGUF version of Qwen3.6-27B for a while, so when Unsloth released an NVFP4 version I wanted to see what it could do in vLLM. Just to avoid confusion: this is Unsloth's Qwen3.6-27B NVFP4 release, not NVIDIA's separate NVFP4 release. The main thing I wanted to figure out was how to set num_speculative