Skip to content
r/LocalLLaMA · Communities

DS V4 on single b300. only 770 tok/s batched in vLLM

Been running DeepSeek-V4-Flash for an offline batch job (cleaning a big pile of short text records, so lots of small prompts rather than chat). Single B300, vLLM 0.25.0, in-process LLM.chat over the batch. Reasoning on, roughly 300 output tokens per item. Best I cn get so far is about 770 aggregate output tok/s at batc