DeepSeek-V4-Flash-0731 UD-IQ3_S 12.5 tok/s on RTX 3090 +128GB DDR5
I managed to run DeepSeek-V4-Flash-0731 UD-IQ3_S in text-generation-webui with: RTX 3090 24 GB 128 GB DDR5 overclocked to 5600 MHz using AMD EXPO llama.cpp loader First,…
I managed to run DeepSeek-V4-Flash-0731 UD-IQ3_S in text-generation-webui with: RTX 3090 24 GB 128 GB DDR5 overclocked to 5600 MHz using AMD EXPO llama.cpp loader First,…
Hello fellow local AI people! I took "you must create your own benchmarks" literally, and built a website for this. How does the end result look…
Been using Qwen 3.6 35B-A3B quite extensively lately and honestly, I’m pretty happy with it. Also tried a few community improvements like Ornith 1.0, which add…
DeepSeek V4 Flash got me thinking... We keep seeing smaller models get way better. A model at a certain parameter count today can be much smarter…
This pull request was added to the main llama cpp about 12 hours ago. I was experiencing some looping and poor behavior yesterday but haven't had…
I'm building a render monster for work, and the idea is to stuff as much GPU power in there as possible. GPUs are going to be…
Same harness, same task set, same agent scaffold, the only thing I swapped was the executor. Not a proper benchmark, no clean tok/s numbers, this is…
The biggest issue with preview was its inability to follow rules prompts and skills. It seems like no matter what you do it ignores them. I've…
For anyone interested, here are the llama-bench results on 3 bit K_XL quantization. I think this could be pushed further but no luck so far. Command…
https://preview.redd.it/u3xo0qklksgh1.png?width=680&format=png&auto=webp&s=9d72c80bbad2559e210897e137d30c88faa7bbc7 Hehe. submitted by /u/laterbreh [link] [comments]