Skip to content
r/LocalLLaMA · Communities

I merged fixes for quantized KV cache into my DeepSeek V4 branch

Check it out: https://github.com/fairydreaming/llama.cpp/tree/dsv4 They are PRs #25247, #25303 (mine) and #25202 (from am17an) but I omitted some padding changes from the last one that I think are not necessary. So if it crashes for you let me know. Also some perplexity values: f16: $ ./bin/llama-perplexity -m ~/ggufs/