r/LocalLLaMA
· Communities
Deepseek v4 Flash on 80 GB VRAM and 128 GB DDR4 RAM
I am using unsloth Q8 Deepseek v4 Flash. So far I am able to run properly with the following command CUDA_DEVICE_ORDER=PCI_BUS_ID CUDA_VISIBLE_DEVICES=0,2,1 llamacpp/llama.cpp/build/bin/llama-server --model unsloth/DeepSeek-V4-Flash-GGUF/UD-Q8_K_XL/DeepSeek-V4-Flash-UD-Q8_K_XL-00001-of-00005.gguf --port 8001 --al