r/LocalLLaMA
· Communities
BeeLlama.cpp v0.4.0: KVarN, KV precision tail, q2_0-q3_1 KV cache, upstream rebase
TL;DR llama.cpp fork with more KV cache quantization features, with all claims supported by benchmarks: KVarN, KV cache precision tail, additional types of standard KV cache (q2_0-q3_1, q6_0, q6_1), and more. BeeLLama v0.4.0 is a complicated update, removing and adding features in roughly equal proportions. Previously