r/LocalLLaMA
· Communities
What is the meta for running Qwen 3.6 27B at 5-10 tok/s as cheap as possible (without speculative decoding)?
The reason I exclude speculative decoding is because I plan on using like DFlash or DSpark for Qwen 3.6, so 5-10 forward passes a second is what I am asking for really. submitted by /u/Aggravating-Push-207 [link] [comments]