You really should not quantize KV Cache for DeepSeek V4 Flash
I don't think anyone should quantize the KV with DS4F. I checked the the quality impact (PPL, KLD, Same TopP) for swhitching from BF16 KV to…
I don't think anyone should quantize the KV with DS4F. I checked the the quality impact (PPL, KLD, Same TopP) for swhitching from BF16 KV to…
M1 Ultra 128GB, Unsloth UD-IQ3_XXS, wired limit at 120GB. I was at 5-6 tok/s before the patch. Getting 15-16 tok/s now with the patched engine, and…
TensorSharp supports DSpark on Deepseek v4 Flash 0731 now. Here is the benchmark result on 4x Nvidia A40 GPUs, cuda 12.8 with/without DSpark: Model: DeepSeek-V4-Flash-0731-UD-Q8_K_XL from…
You have two choices here (in order of pref): Downgrade CUDA from 13.3 to 13.1 (skip 13.2 due to bugs)
submitted by /u/rmhubbert [link] [comments]
About the new deepseek v4 flash version / update, does anybody now about the new values about: MMLU-Pro GPQA Diamond TruthfulQA About the other values, its…
> After using DeepSeek v4 Flash for a day, its Harness is much better than Kimi K3!!!IntriguingYufan Sheng: 用了一天 DeepSeek v4 Flash,它的 Harness 比 Kimi K3…
There had been no “preview subsidy”SiliconFlow prices are distinct from DeepSeek’s and in particular worse on cacheNevertheless they do work together and this is a minor…
I spent a while debugging my local DeepSeek V4 Flash setup and wanted to share a few lessons from the process in case it saves someone…
Wenfeng is too cultured to really act like this, but he totally feels this wayDeepSeek people have a lot of pride in what they doGorden Sun:…
The biggest issue with preview was its inability to follow rules prompts and skills. It seems like no matter what you do it ignores them. I've…
Thanks to the community help I finally launched this llm. LM Studio refused to load weight onto second GPU but Unsloth Studio did so everything was…
submitted by /u/perelmanych [link] [comments]
The model refuses to load into VRAM and uses only RAM. What can be an issue? Q2_K_XL from Unsloth if that changes something. submitted by /u/esw123…
RT cheatyhelp me understand the cope deepseek doomers are going through when they can't even write a tweet on their ownwhy are you making shit up…
the bots are precisely wrong, this is the opposite of DeepSeek's speciality. But come to think of it, DeepSeek is very Tau Law pilled. What they…
What speeds are everyone getting with deepseek v4 flash 0731? I’m getting~200 tps prompt processing / ~11 tps token gen, on 4x5060ti16gb with ddr4 3200 ram…
RT Andrew HoThe ability for DeepSeek to be price competitive with GPT-5.6 Luna is genuinely incredible. The inference team at OpenAI was brilliant from what I…
I'm an avid user of Deepseek v4 Flash via antirez's DS4 DwarfStar inference engine, and so when the new checkpoint dropped, the first thing I did…
DeepSeek published DeepSeek-V4-Flash-0731 on Hugging Face and moved the official V4-Flash API into public beta on July 31, 2026. The model card is explicit that this…
https://preview.redd.it/h7zv5tb3tmgh1.png?width=2854&format=png&auto=webp&s=507380e8f862c18f10f7c5c84da9e8d1c59139b0 Deepseek's new flash model is unexpectedly cheap and high-performing across useful benchmarks. It's priced at $0.09 / $0.18 per 1M. Truly "intelligence too cheap to…
submitted by /u/InternationalGap3698 [link] [comments]
I was surprised to see that Deepseek V4 Flash is extremely smart and small enough to fit in setup that can be built with < $50,000.…
submitted by /u/curiousily_ [link] [comments]
DeepSeek shocked the field in late 2024 with DeepSeek-V3, then again with R1, a reasoning model trained at a fraction of the budget Western labs spend. Current lines: V3.2-Exp, R1, V3. The most-talked-about Chinese AI lab of the cycle.
Owner: DeepSeek. We have 280 stories indexed for this model, auto-tagged from titles across every tracked source — official announcements, papers, GitHub release notes, and third-party press. The CTA on each card links to the original; the official site is www.deepseek.com.
Related text models: GPT, Claude, Gemini, Gemma, Llama, Mistral, Grok, Qwen.