r/LocalLLaMA
· Communities
NInfer RTX 4090 for Qwen 3.8 27B update – up to 250-350K tokens context in VRAM
I've made some improvements to my fork of NInfer, adding rk2v4-e8 quant option for the KV cache, which can reach up to 250-350K tokens of context window depending on the configuration like vision, MTP, etc, on my single RTX 4090 without spilling over into system RAM. On lower context window runs, I also made some optim