r/LocalLLaMA
· Communities
I ran DeepSeek V4 Flash 284B + DSpark on one RTX PRO 6000. The drafter was faster in RAM than VRAM.
Hey guys, Just finished benchmarking DeepSeek V4 Flash 284B + DSpark on a single RTX PRO 6000 96GB. Short version: DSpark: ~15–17% faster generation on my coding workload On this setup, the DSpark drafter was faster in system RAM than VRAM q8_0 KV cache: 256K → 768K context with basically no decode-speed loss Best 9-tu