r/LocalLLaMA
· Communities
DFlash made Laguna S 2.1 (71 GB Q4) 2.5x slower on 2x RTX 5090. I tuned it from 23 to 64 tok/s, benchmarked on Spec-Bench, and I’m still running without it
I ran Laguna S 2.1 (118B MoE, 71 GB Q4) with DFlash speculative decoding on 2× RTX 5090. It doesn't fit, experts spill to CPU RAM. • default flags: 23 tok/s vs 58 without the draft. 2.5× SLOWER • tuned: 64 tok/s vs 62 baseline • even ONE 5090: 33 vs 30. Actually usable. https://preview.redd.it/gqkq9ceqwzeh1.png?width=3