r/LocalLLaMA
· Communities
nvfp4 kv-cache on 2×5060 ti, vllm
I can't take credit for this, someone else described the approach, and I just copy/pasted it into opencode/GLM 5.2 to get it working. It appears to be working: https://github.com/vllm-project/vllm/issues/49011 https://preview.redd.it/k1whe1qrzgeh1.png?width=2446&format=png&auto=webp&s=8886c34ec23d1949bcafaa570dfb8edd08