60-82% accuracy swing on 4B model classification task: the only variable was harness design
I ran a pre-registered ablation on a classification task (Kubernetes issue → SIG triage) using a 4B model on a 6GB laptop GPU. Same frozen weights,…
I ran a pre-registered ablation on a classification task (Kubernetes issue → SIG triage) using a 4B model on a 6GB laptop GPU. Same frozen weights,…
I can run the mediums, but sometimes I want a faster option that's smarter than Qwen 27B/35B. On my hardware I get like 500 to 800…
https://preview.redd.it/h7zv5tb3tmgh1.png?width=2854&format=png&auto=webp&s=507380e8f862c18f10f7c5c84da9e8d1c59139b0 Deepseek's new flash model is unexpectedly cheap and high-performing across useful benchmarks. It's priced at $0.09 / $0.18 per 1M. Truly "intelligence too cheap to…
Its there now! Edit: They just posted it - Q1, Q2, and Q3. Time to rock and roll submitted by /u/live4evrr [link] [comments]
Hello guys, I'm curious about running DeepSeek-V4-Flash-0731 locally. Since it’s a Mixture of Experts (MoE) model with only 13B active parameters, I was hoping the VRAM…
submitted by /u/InternationalGap3698 [link] [comments]
I was surprised to see that Deepseek V4 Flash is extremely smart and small enough to fit in setup that can be built with < $50,000.…
submitted by /u/curiousily_ [link] [comments]
4-bit - 155gb 8-bit - 162gb "Smaller ones are coming" submitted by /u/RunawayPeeko [link] [comments]
Source: https://x.com/deepseek_ai/status/2083084415157022911 & https://deepswe.datacurve.ai/ just combined data view. DeepSeek claims, not verified by DeepSWE yet. submitted by /u/sdexca [link] [comments]