Expert-only IQ3 requant of DeepSeek-V4-Flash-0731: better KLD than UD-IQ3_S, 1.4x decode on a CPU-spill rig
Hey all, tldr / who this helps: you run a mixed multi-GPU box where the experts spill to RAM, and you want to stay in the…
Hey all, tldr / who this helps: you run a mixed multi-GPU box where the experts spill to RAM, and you want to stay in the…
I know in in this community LLM's are generally used for coding but there are other usecases besides coding and those usecases should be tested too.…
Hi everyone, I’m fairly new to the multi-GPU side of local LLMs and I’m trying to understand how inference actually scales across multiple GPUs. Suppose I…
https://preview.redd.it/6aodizfzpugh1.png?width=1252&format=png&auto=webp&s=5749d622a050de8c325093549bb0cd31b1050cd9 I am trying to learn Japanese so I vibe coded scripts that convert book page images into a website with Kokoro TTS voiceovers and contextual…
Hello guys, hoping you're doing fine. Lately with all the new models, and how popular is offloading, what are your min good or usable t/s for…
submitted by /u/Fcking_Chuck [link] [comments]
Here are my numbers: Quant Size Layout Decode Prefill Draft acceptance UD-Q8_K_XL 150.8 GiB 20 layers CUDA0 / 23 ROCm0 + drafter 44.0 t/s 564 t/s…
I managed to run DeepSeek-V4-Flash-0731 UD-IQ3_S in text-generation-webui with: RTX 3090 24 GB 128 GB DDR5 overclocked to 5600 MHz using AMD EXPO llama.cpp loader First,…
Hello fellow local AI people! I took "you must create your own benchmarks" literally, and built a website for this. How does the end result look…
Been using Qwen 3.6 35B-A3B quite extensively lately and honestly, I’m pretty happy with it. Also tried a few community improvements like Ornith 1.0, which add…