r/LocalLLaMA
· Communities
Optimised DSv4-Flash for 2x GH200: 10,000 tok/s PP, >300 tok/s TG on SGLang
There are some PRs to use and a nice trick to speed up PP on really longs contexts in my write up. Hope it helps! TL;DR: On this dual GH200 box, you build vLLM v0.26.0 from source, add the merged DSV4 cache-layout patch (PR #48993), disable async scheduling, and run DSpark at 6 predicted tokens to give: ~276 decode tok