r/LocalLLaMA
· Communities
DeepSeek v4 Flash on 5090 in llama.cpp with 1 Million context
After the recent llama.cpp changes, DeepSeek V4 Flash has become much more usable. I ran some benchmarks and wanted to share the results along with the config I used. I'm using DeepSeek-V4-Flash-UD-Q8_K_XL from Unsloth: https://huggingface.co/unsloth/DeepSeek-V4-Flash-GGUF Config: llama-server -m DeepSeek-V4-Flash-UD