Skip to content
r/LocalLLaMA · Communities

Qwen3.6 35B A3B KV cavhe quantizations memory footprint

Is it really worth it to quantize KV cache below Q8 accepting heavy trade-off submitted by /u/token---- [link] [comments]