X · @teortaxesTex
· X / Twitter
> 464 tok/s on batch size 1 with 4×4 GB300 but I don't care about batch size 1…
> 464 tok/s on batch size 1 with 4×4 GB300but I don't care about batch size 1…vLLM: vLLM hit new peak bs=1 decode on Kimi-K3: 464 tok/s 🚀Under a low-entropy reasoning workload, Kimi-K3 + DSpark on vLLM reaches 464 tok/s on batch size 1 with 4×4 GB300.This benchmark is fully reproducible with public image: vllm/vllm-ope