Skip to content
r/LocalLLaMA · Communities

Qwen3.8-27b on RTX 3090 – 82 tps single request, up to 672 tps peak

Hi, After a long night of optimizations, I believe I have made the fastest inference engine for Qwen3.6-28B on a 3090. Quick metrics: - 250w power capped - Up to 195k context (ships with 150k for safety though) - 82 tps single request, 417 tps sustained with 64 concurrent - Between 17% to 149% faster than ninfer depend