Skip to content
r/LocalLLaMA · Communities

Benchmark – 4x 5060 Ti (64GB VRAM) (P2P) – Qwen3.6 27B (INT8 /w bf16 kv cache) @ 8 concurrency with SGLang. SGLang seems to handle higher concurrency better with this setup

I recently posted some posts with VLLM showing issues with TTFT and concurrency with 4x 5060 ti's. Wanted to share this benchmark to provide what worked for me so other people that are planning to go the 4x 5060 ti route aren't discouraged. Benchmark Results ============ Serving Benchmark Result ============ Backend: s