Skip to content
r/LocalLLaMA · Communities

880 tok/s on one 5090 Qwen3.8-27B in 4-bit NVFP4, full 262k context

Numbers first, on a single RTX 5090, Running CachyOS with COSMIC, and the entire desktop costs about 150 MB of VRAM. 880 tok/s aggregate at 6 parallel requests (peaked at 967 on one run). 200+ tok/s single stream with MTP speculative decoding ~5,950 tok/s prefill (llama.cpp Unsloth Q5_K_XL manages ~1,700 on the same bo