r/LocalLLaMA
· Communities
DeepSeek-V4-Flash-0731-UD-Q3_K_XL 3×3090 test results
For anyone interested, here are the llama-bench results on 3 bit K_XL quantization. I think this could be pushed further but no luck so far. Command ./llama-bench -m /home/user/llamacpp/modelsmain/unsloth/ds4/DeepSeek-V4-Flash-0731-UD-Q3_K_XL-00001-of-00004.gguf -ngl 21 --split-mode layer -p 512 -n 128 -r 5 Output CUDA