DeepSeek-V4-Flash-0731 is going to cause another market crash.
Beats GLM 5.2, and is the same cost as the previous one. submitted by /u/Potential_Top_4669 [link] [comments]
Beats GLM 5.2, and is the same cost as the previous one. submitted by /u/Potential_Top_4669 [link] [comments]
openPangu-2.0-Pro is an MoE model trained on Ascend. The model has 505B total parameters and 18B activated parameters. Its context length is 512k. The total pretraining…
Deepseek's new model V4 Flash 0731 is much better, I (Claude lol) did a bit of linear regression with a leave one out style verification to…
DeepSeek V4 Flash: Preview → 2026-07-31 Benchmark Preview 0731 Δ Terminal Bench* 56.9 82.7 +25.8 Toolathlon 51.8 70.3 +18.5 NL2Repo — 54.2 new Cybergym — 76.7…
submitted by /u/mineyevfan [link] [comments]
https://api-docs.deepseek.com/updates/ submitted by /u/Nunki08 [link] [comments]
So now full 512gb of vram, I think I will need to get a third one soon seeing that llms keep getting absurdly huge! submitted by…
I've ran Kimi-k3 through 34 oneshot prompts and evaluated the generated htmls, screenshots and gifs using sonnet 4.6. It came out to be better than opus4.8…
This post was not written by a clanker. Hey guys, I'm a comp sci major who wanted to introduce a cool project I built for quantizing…
https://x.com/MiniMax_AI/status/2083006198828417501?s=20 Quote from their article: Today, we're launching MiniMax H3, a general-purpose multimodal generation model. H3 understands unified context across text, images, video, and audio, generating…