Skip to content
r/LocalLLaMA · Communities

Spiritbuun’s VBR (Variable Bit Rate) KV cache — first impressions

Well, this is an appreciation post on the spiritbuun llama.cpp fork. Some weeks ago I was testing forks and configurations to find which one was the best for my secondary model on my 3060, and the winning combo turned out to be Spiritbuun's fork + CUDA + mudler's Apex I-Compact quantization for the Qwen3.6-35B-A3B mode