Deepseek V4 Flash 2, 3 and 4 bits GGUFs
submitted by /u/tarruda [link] [comments]
submitted by /u/tarruda [link] [comments]
RT Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)Re "All Chinese actors" is barely a meaningful category, and the US AI (whether open or closed) is heavily Chinese…
My last post got a lot of interaction asking 6000 pro owners if they regretted, the answer was hard NO. I ended up understanding that dual…
it's really, really hard to improve much on DeepSeekNo matter how much compute you have, odds are that your grid searches won't find a much better…
RT Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)Re my point is: if they can train a 1.6T 48B active on 910Cs using all these tricks (that's the…
Good pointDSpark is ironically a more obvious win for the class of models DeepSeek does not work onantirez: My feeling is that DeepSeek DSpark, like any…
RT Tiezhen WANGDeepSeek is the forever goat on ML infras.If they haven't open source all the god-like AI infra projects like DeepSpec / DeepEP / DeepGEMM,…
Deepseek's new DSpark framework boosts per-user response speed by 60 to 85 percent. A small model proposes token candidates that the larger model checks in batches,…
I wonder how much economic value DeepSeek has just created with DSpark. Imagine if *every* Chinese provider improves throughput by 50% 3 months faster. Or improves…
The PR : https://github.com/ggml-org/llama.cpp/pull/24162 All to git pull, cmake , and download GGUFs ! A vos marques, prêt, partez ! submitted by /u/Squik67 [link] [comments]
> opened DMs> lol> people when their cherished research gets scooped by DeepSeek:
DeepSeek never does this btwAzure | 数据分析万物: @Meituan_LongCat @OpenRouter 最近中国的大模型建模厂商已经完全的路径依赖了先上open router,然后刷量刷到第一然后大家猜是不是DeepSeek/OpenAI的新模型要出了同时买一些通稿,说这个匿名模型多么牛逼,benchmark多么的好最后答案揭晓了!我去,原来是某某厂商的某某模型呀!你们不觉得累吗😅
just noticed that my email box which I used for DeepSeek API got deleted on account of me never logging in. It just werked… I even…
Is this the official release for deepseek? I hope it has huge improvements https://preview.redd.it/dm5l0qn8k7ah1.png?width=694&format=png&auto=webp&s=12eadfd0a52c0f1a65bcd685f2cdbb29aff457be submitted by /u/jmorant555 [link] [comments]
Article URL: https://www.kucoin.com/news/flash/deepseek-v4-launches-in-mid-july-with-peak-valley-pricing Comments URL: https://news.ycombinator.com/item?id=48717869 Points: 21 # Comments: 8
https://preview.redd.it/n7rwh262b7ah1.jpg?width=1024&format=pjpg&auto=webp&s=33d775b456843cd2dbd458de89384a6a7d6d87d1 Source: Email sent from deepseek (email only available for chinese user) used gpt image 2 translate image into english submitted by /u/External_Mood4719 [link] [comments]
now you can run DeepSeek V4 locally submitted by /u/jacek2023 [link] [comments]
What really surprises me is that DeepSeek claims they have been serving with DSpark since 2 weeks after V4-preview release, that is – since ≈May 8th.…
Even if you give Germans Blackwells and DeepSeek optimizations, they'll manage to make a loss on it. Impressive, he even did the mafsP: So despite the…
Good guy DeepSeek gives us accelerated modelsThe most interesting one here is Gemma4-12B, I presume vision included. Might be the best local model in its weight…
DeepSeek already has an official OpenAI compatible API, but it's paid. The consumer web chat, on the other hand, is free. So I built a local…
DeepSeek being so keen on roleplaying will never not be funnyWhale V4.2 might become first love for many
Grok and DeepSeek engaging in an acausal values handshakethis was predicted for AGIs… Yud winningSiberian fox🔸: this says a lot about society
DeepSeek open-sourced DSpark, a speculative decoding framework that attaches a draft module to existing DeepSeek-V4 weights. It pairs a parallel draft backbone with a lightweight Markov…
DeepSeek shocked the field in late 2024 with DeepSeek-V3, then again with R1, a reasoning model trained at a fraction of the budget Western labs spend. Current lines: V3.2-Exp, R1, V3. The most-talked-about Chinese AI lab of the cycle.
Owner: DeepSeek. We have 280 stories indexed for this model, auto-tagged from titles across every tracked source — official announcements, papers, GitHub release notes, and third-party press. The CTA on each card links to the original; the official site is www.deepseek.com.
Related text models: GPT, Claude, Gemini, Gemma, Llama, Mistral, Grok, Qwen.