Skip to content
r/LocalLLaMA · Communities

DeepSeek V4 Flash | IQ3_XXS-AS & IQ2_S Bench | mainline b10064 vs fairydreaming | 1xRTX 3090 + 128GB DDR4 | 250PP/11TG on 50K CTX

Hey all! Wanted to see how DeepSeek V4 Flash GGUFs in two different quants perform on my hardware and share the results. Tested two quants on the fairydreaming/llama.cpp dsv4 fork. As a bonus, I also ran the same model (IQ3_XXS-AS) on mainline llama.cpp b10064 just to check. TLDR: turns out the fork is pointless now! M