r/LocalLLaMA
· Communities
DKV: Open-source KV-cache compression framework for local LLM inference (CLI + technical report)
Hi everyone! Over the past five months I've been working on DKV (DifferentialKV), an open-source project exploring KV-cache compression for long-context local LLM inference. The goal is to reduce KV-cache memory requirements through anchor-based representations, joint low-rank compression, exact residual preservation,