r/LocalLLaMA
· Communities
Concurrency plus nvfp4 on Blackwell
Parsed from VLLM log file ~2000 tps in aggregate performing bulk captioning on images. Above is parsed from vllm log while a client runs 30 concurrent streams, each concurrent stream has 1 request with an image and prompt, then a 2nd request on the same stream (so 1st Q:A would be cached). Typical log line: Engine 000: