Skip to content
r/LocalLLaMA · Communities

NInfer RTX 4090 for Qwen 3.8 27B update – up to 250-350K tokens context in VRAM

I've made some improvements to my fork of NInfer, adding rk2v4-e8 quant option for the KV cache, which can reach up to 250-350K tokens of context window depending on the configuration like vision, MTP, etc, on my single RTX 4090 without spilling over into system RAM. On lower context window runs, I also made some optim