Skip to content
r/LocalLLaMA · Communities

Kindly Benchmark Higher Quants of DeepSeek-v4-flash Against Qwen-3.6-27B Q8!

Kindly Benchmark Higher Quants of DeepSeek-v4-flash Against Qwen-3.6-27B Q8! I am running the UD-Q2_K_M of the model locally, though I can run Qwen3.6-27B_Q8_K_XL at around 70t/s with MTP activated. The question I am constantly asking myself is: Is it worth running a slower higher quantized version of the Deepseek-v4-f