Skip to content
r/LocalLLaMA · Communities

Ling 3.0 Flash on Strix Halo

vLLM ROCm/HiP, 4 bit compressed-tensors (int4) Not a fair comparison, but Qwen-122b on the most optimized format possible I have run (rocmFP4) does not touch Ling in speed. https://x.com/ciruai/status/2085996633267777554?s=46 Tool call is broken in certain harnesses. It works well with pi-type harnesses (omp, feynman).