Deepseek V4 Flash just hit Colibri, does anyone have numbers?
I'm mosty interested in 128-192GB VRAM with 128-256GB RAM to spare, so SSD streaming is basically not even necessary. Seems only FP4 is supported, so older…
I'm mosty interested in 128-192GB VRAM with 128-256GB RAM to spare, so SSD streaming is basically not even necessary. Seems only FP4 is supported, so older…
Inspired by a post from u/giveen I motivated claude (no patinence on my side to work through everything myself) to help me get DS running on…
🤯 DeepSeek reports V4 Flash at 82.7 on Terminal-Bench 2.1, ahead of V4 Pro Preview at 72.1, despite using roughly one-fifth the total parameters.
TensorSharp's MoE CPU-offload feature has been merged into main. Here is the parameters description of this feature: Mixture-of-Experts CPU offload: --n-cpu-moe | -ncmoe Keep the routed…
RT PiDeepSeek V4 Flash is Ollama's fastest growing model ever in token usage, and the most popular model on OpenRouter this week.It’s available in Pi across…
It's the purple cluster on the top left (the good corner...) I'm running the MXFP4 version from Bartoswski with Dspark at 1K t/s prefill and 90…
Alibaba has launched Qwen3.8-Max, its largest AI model to date, as DeepSeek’s latest V4-Flash model draws attention for inference pricing that is lower than several competing…
So one of yall mentioned that cuda 13.1 or 13.2 is broken for unsloth so I looked in to it, and they were right. I had…
Just wanted to share my agentic coding benchmark run of DSv4F 0731 at both High and Low reasoning efforts (not Max)... I ran a 109-question subset…
https://preview.redd.it/522fsdwvtdhh1.png?width=1200&format=png&auto=webp&s=6a6cf7a467514167a8193029dbd20fb3a9ba4f6c It ranks lower than both Sonnet 4.6 and Luna. I'd wager Luna costs in the same ballpark as DS4F considering Luna’s token efficiency. DeepSeek being…
I've been working on speeding up DeepSeek-V4-Flash-0731 in Krasis and have now got the long-prompt prefill quite a bit faster on a single RTX PRO 6000…
I really like to use this one SQL benchmark when testing new models. I had another post some time ago with my benchmarks, but I decided…
submitted by /u/tarruda [link] [comments]
Currently running some experiments using the streamed experts trick that's been floating around this sub as well as some of my own trickery to get prefill…
Article URL: https://github.com/ryanzhou/deepseek-v4-flash-mi300x Comments URL: https://news.ycombinator.com/item?id=49166386 Points: 29 # Comments: 4
RT Andrew CurranAt first I thought Bloomberg forgot to add DeepSeek's pricing to the chart.
> While the new DeepSeek V4 Flash performed similarly to the preview checkpoint in our agentic coding harness, it's dramatically stronger in one-shot, fluid intelligenceUnexpected but…
Ran DeepSeek V4-Flash-0731 — the full official checkpoint, not a re-quant — on commodity used hardware. Sharing because I couldn't find anyone else publishing Ampere results…
https://preview.redd.it/uybjxyypj7hh1.png?width=1200&format=png&auto=webp&s=8293e8da332a14920b335caee52473762c09530d We just tested the big 3 of open-weight models on our internal evals. The tasks include long-running workflows involving multiple applications (Pagerduty, Gmail, HubSpot, Airtable,…
Has anyone figured out how to enable speculative decoding with deepseek v4 flash 0731 on llamacpp? I’m on the right release for llamacpp (b10228 or earlier)…
Qwen3.8-Max (2.4T) is another massive contribution to the open weight community. On benchmarks, it performs closely to Kimi K3 and DeepSeek V4 flash across all categories…
Continuing my weekend of oneshotting the cheap OpenRouter models, here are all 10 DeepSeek models across the same 35 prompts. DeepSeek had a rougher time (more…
Originally, I was only getting around 140pp/s and about 21tg/s, but the config with -b 8192 -ub 8192 --cpu-moe is vastly superior, let's say 700pp/s and…
Id figured since they first emailed people about api price changes coming mid july then delayed the v4 flash release to late july, I wonder if…
DeepSeek shocked the field in late 2024 with DeepSeek-V3, then again with R1, a reasoning model trained at a fraction of the budget Western labs spend. Current lines: V3.2-Exp, R1, V3. The most-talked-about Chinese AI lab of the cycle.
Owner: DeepSeek. We have 280 stories indexed for this model, auto-tagged from titles across every tracked source — official announcements, papers, GitHub release notes, and third-party press. The CTA on each card links to the original; the official site is www.deepseek.com.
Related text models: GPT, Claude, Gemini, Gemma, Llama, Mistral, Grok, Qwen.