Skip to content
r/LocalLLaMA · Communities

Comparing local inference speeds across a few real setups people are running (3090 vs 5090 vs dual 6000)

Pulled together token rates from a few different local rigs people have reported running lately, just to get a sense of what's realistic at each hardware tier(source discord group) Qwen3.6 27B on a single 3090 (Q4/Q8 MTP, 128k ctx): ~50 tok/s inference, ~950 tok/s prompt processing Qwen3.6 27B on a 5090 (Q6 MTP, tuned