The official release Deepseek V4 flash is live on the API
submitted by /u/mineyevfan [link] [comments]
submitted by /u/mineyevfan [link] [comments]
I've noticed that there's some confusion about Luna vs. DeepSeek pricing, so I've made a table. Luna does win on extremely cheap input tokens. But in…
I can't replicate this, and maybe this is "routing to Fable" or whatever. But in theory, DeepSeek could continue the pretraining of V4s. They have intermediate…
The reason Luna blows DeepSeek out of the water is that it's just stronger than V4-Pro.It's not cheaper on a token basis in agentic sessions. Cache…
We recently used DeepSeek V4 Flash as a teacher for finance tasks with GPT-OSS-120B. Distillation works well on this problem. At a constrained 8k token budget,…
Hey fellow llamas, sorry for posting again this week but i thought this was interesting to showcase to share with y'all. Lucebox partnered up with AMD…
I feel more confidence in DeepSeek's "we will delay V4 until we co-optimize the model and the harness enough" decision now. Cloning Claude Code won't cut…
Hey fellow llamas. we have something new for Strix Halo owners we thought would be useful to share. i'll keep it short: We were able to…
Confirming the obvious: Kimi K3 has the best kv cache economics out of all major models except DeepSeek V4 (and even then, it’s almost the same…
> (Maybe)> nothing like this is confirmed at allI guess that recent confusing behavior of DeepSeek was about harness experiments, and they really didn’t want Fable…
Dreadful thought:- Wenfeng says that half of core researchers, his "most important people", are doing data annotation- DeepSeek is rapidly expanding by >100%, which means core…
This is much more retarded than the DeepSeek Moment selloff. Chinese DUV won't threaten ASML's slice, first of all because Chinese (strategically driven) demand for it…
DeepSeek has hit a wall (on WeirdML)no progress since 3.2-Specialeskull charted like a wypipoit's over…
Dario opens his newest essay with bullshitting. In "DeepSeek and Export Controls" as well as in "Two Scenarios" and in "Policy on the AI Exponential", he…
So that, too, is the "closed models" that the open models Coalition supports. I've just recently thought: SSI and DeepSeek are surprisingly similar. Sole focus on…
I ran DeepSeek V4 Flash through Claude Code, OpenCode and Pi on my own benchmark, and the quality came out basically the same across all three…
Article URL: https://github.com/demo-zexuan/liang-wenfeng-investor-meeting-2026-7-22/blob/master/%E6%A2%81%E6%96%87%E9%94%8B%E6%8A%95%E8%B5%84%E8%80%85%E4%BA%A4%E6%B5%81%E4%BC%9A-%E6%96%87%E5%AD%97%E7%A8%BF_1_18_translate_20260723201651.pdf Comments URL: https://news.ycombinator.com/item?id=49052912 Points: 21 # Commen
I understand that the laguna model is either still buggy or potentially benchmaxxed. So I’d like to know for people who really tested, are DS flash…
Hallo Everyone, i need to generate big amount of high quality data, for that i need some cheap API providers. who is the cheapest,most reliable provider…
On the contrary, full vertical integration with proprietary hardware is when you can open source with impunity. I think that's what DeepSeek will do. They'll codesign…
How is Flo even real? What is this discourse for The Raped?the CCP won't be harmed by you dropping DeepSeek…Just use ChatGPT and uh, Nemotron, lil…
Three years and one month ago, Mustafa Suleyman announced having a cluster that's larger than anything DeepSeek had at the point of releasing V4-Pro. Roughly two…
meanwhile: «The China National AI Industry Investment Fund invested in DeepSeek and received voting rights. The fund has committed RMB $1B.»I suspect Wenfeng can be a…
TLDR: I (with the help of AI) re-implemented every Blackwell-only kernel (DeepGEMM, FlashInfer sparse-MLA, block-scaled FP8) in Triton, because they simply don't exist for sm89. The…
DeepSeek shocked the field in late 2024 with DeepSeek-V3, then again with R1, a reasoning model trained at a fraction of the budget Western labs spend. Current lines: V3.2-Exp, R1, V3. The most-talked-about Chinese AI lab of the cycle.
Owner: DeepSeek. We have 280 stories indexed for this model, auto-tagged from titles across every tracked source — official announcements, papers, GitHub release notes, and third-party press. The CTA on each card links to the original; the official site is www.deepseek.com.
Related text models: GPT, Claude, Gemini, Gemma, Llama, Mistral, Grok, Qwen.