r/LocalLLaMA
· Communities
It’s the small things that matter the most. – llama.cpp – Bunch of updates(Boost & Fixes)
ggml-cuda: add chunked SSD matmul for Mamba-2 prefill acceleration- #22675 Nemotron-Nano-9B-v2 ub base (scan) branch (SSD) speedup 128 5,404 5,351 −1% (both scan) 256 6,180 7,110 +15% 512 6,627 7,778 +17% 1k 6,814 8,152 +20% 2k 6,759 8,190 +21% 4k 6,660 8,118 +22% 8k 6,387 7,761 +22% pp16384 tok/s (base=scan, branch=SS