r/LocalLLaMA
· Communities
I benched quad 20GB 3080s on Vast AI for code generation with Qwen3.6-27B so you don’t have to (it’s even better than quad 5060Tis)
TL;DR it's pretty goddamned fast; 69 tps decode at near max (256k) context with MTP on. prefill numbers went down to 893 at max context with prompt cache turned off. https://jdkruzr.github.io/3080bench/ here's how the tests were run: https://github.com/jdkruzr/3080bench/ there is probably more performance left on the t