Deepseek V4 Flash running on RTX 5090 MoE
Here is the results of optimizing it for my setup: Benchmark results of the optimisation showing TG T/S from 22.7 to 21.3, and PP T/S from…
Here is the results of optimizing it for my setup: Benchmark results of the optimisation showing TG T/S from 22.7 to 21.3, and PP T/S from…
Hey all, So I tend to favor the Claude Desktop app in Code mode as the GUI does a great job of previewing code, MCP browser…
submitted by /u/xw1y [link] [comments]
https://github.com/IceFog72/llama.cpp I added an experimental sampler to llama.cpp called scatter. The short version: it slightly smooths the model’s next-token probability distribution inside the already-selected top candidates.…
I run it on i5 6500 and I get 9t/s its really fast and the output is a lot better than ChatGPT 3.5 and maybe its…
Dspark looks very exciting. Anybody got insight into whether it can be added to qwen 27b? submitted by /u/GotHereLateNameTaken [link] [comments]
I just heard about SwiReasoning and tried it out on Qwen 3.6 27b and im kinda surprised. Its answers are more on point and it solves…
The reason I exclude speculative decoding is because I plan on using like DFlash or DSpark for Qwen 3.6, so 5-10 forward passes a second is…
So here are some docs for getting API Keys for them, because Google loves to show Reddit posts: https://developer.ant-ling.com/en/docs/models/ring/ https://longcat.chat/platform/docs/ For longcat I had to go…
- 2x RTX Pro 6000 Max-Q (96GB) - 8x RTX 3090 (24GB) - 2x RTX 5090 (32GB) - 3 PSUs - 128GB DDR5 SDIMM RAM (4-channel)…