r/LocalLLaMA
· Communities
For V100 Users: SGLang running Qwen+Dflash and Laguna
Forked SGLang, wrote TeilLang FlashAttention for V100, used open-source marlin-v100, ungated flashinfer for sm70, made Dflash work for Qwen3.5/3.6 models, added Laguna S2.1 support, tried to make dflash work for Laguna(and no luck so far). ~4000-6000pp, ~100 tks tg(Qwen only). Running on my 4xV100 32GB NVLINK: https://