r/LocalLLaMA
· Communities
Qwen3.6 35B A3B KV cavhe quantizations memory footprint
Is it really worth it to quantize KV cache below Q8 accepting heavy trade-off submitted by /u/token---- [link] [comments]