r/LocalLLaMA
· Communities
DS V4 on single b300. only 770 tok/s batched in vLLM
Been running DeepSeek-V4-Flash for an offline batch job (cleaning a big pile of short text records, so lots of small prompts rather than chat). Single B300, vLLM 0.25.0, in-process LLM.chat over the batch. Reasoning on, roughly 300 output tokens per item. Best I cn get so far is about 770 aggregate output tok/s at batc