Anyone else amped up over Qwen 3.8?
I’ve been using 3.6 27B Q4, and that quant is fast on an M5. The code has been average, but consistently “good enough.” And, after a…
I’ve been using 3.6 27B Q4, and that quant is fast on an M5. The code has been average, but consistently “good enough.” And, after a…
I compared Qwen 35B-A3B MoE against Qwen 27B dense on a series of local coding-maintenance tasks. On my R9700/llama.cpp setup, the MoE model generated about 3.9×…
Hi there, I have a setup with 4x 2080Ti 22GB, but am not able to see much benefits in running models larger than 16-20GB in speed.…
Basically what the title says. For me, it would always crash and burn trying to use tensor split. Apparently, there's some bug where GPU memory gets…
- On b10173 - "state":"loading" 4min54sec. - With this PR and GGML_RPC_LOAD_THREADS 12 - "state":"loading" 1min38sec The PR is close to ready, will need a docs…
Tried to use Beellama , and using the kvarn6 flag, i notice that its in llama-server --help but its not working. I must be doing something…
(I am not a native speaker, written by myself, so please bear with me) I really want to like DeepSeek-V4-Flash-0731. But it has serious flaws that…
submitted by /u/johnnyApplePRNG [link] [comments]
Hey all, I'm serving DSv4Flash 0731 on a cluster of 2x DGX Sparks but am running into constant issues with having almost no RAM (unified memory)…
I run the following on a 5090 and have been okay with its performance, it does most things somewhere 80-100 t/s, though that can slow down…