r/LocalLLaMA
· Communities
Benchmark of the new unsloth/Qwen3.6-27B-NVFP4 on 4x 5060 ti’s with P2P and PP=4 at 1,4,8,12, and 16 concurrency.
Been trying to troubleshoot prefill around when I saw the newer (supposibly faster) nvfp4 quant. For those curious of prefill issues on 4x5060 ti's, hopefully this helps paint a picture for at what concurrency prefill/ttft starts being an issue. ## info: PP=4 (new unsloth NVFP4 quant) Pipeline parallelism=4, 4x 5060 TI