r/LocalLLaMA
· Communities
Q2 DeepSeek V4 Flash on 2x 3080 20GB, 64GB DDR5 | 17 tk/s gen, 270 tk/s prefill
Hey, it's my first time posting here and I thought I'd share my progress on getting Antirez's imatrix Q2 DeepSeek V4 Flash GGUF (86.7 GB) running on my build. I used this llama.cpp fork which fixed the model's output when KV cache is quantised to Q8. Specs: - Ryzen 7 7800X3D - Asus ProArt B850-CREATOR WIFI NEO - 64GB D