Skip to content
llama.cpp releases · Infrastructure

b9911

CUDA: Fuse MMVQ post-scale for NVFP4 (#24481) CUDA: Fuse MMVQ for NVFP4 and BS 1 TODO: Add tests to test-backend-ops (did verify correctness manually for one model) Reorder bias/scale once PRs for NVFP4 are merged/landed Add dense MMVQ fusion as well Perf numbers on B4500. Note qwen35 is FP8->Q8 ./scripts/compare-llama