r/LocalLLaMA
· Communities
Benchmarked every spec-decode method on Qwen3.6-27B across vLLM and SGLang (single RTX PRO 6000 Max-Q)
https://preview.redd.it/wluwwp6s4heh1.png?width=1248&format=png&auto=webp&s=6e95d963645c5a0ef750bf81324a6d4dcbc0389e Spent the last few days measuring speculative decoding on Qwen3.6-27B (dense, NVFP4) on one RTX PRO 6000 Max-Q, comparing vLLM and SGLang across MTP, DFlash, EAGLE3 and ngram. Same pinned client for ever