r/LocalLLaMA
· Communities
Has anyone gotten Llama.cpp (or other) working using Intel iGPU (arrowlake) where it actually improves anything?
I Recently did a bunch of tests and wrote them all up on here, but the short version is that Vulkan basically doesn't work (or when it does, it's at 1tok/s at best). SYCL works pretty well, seems to run the Qwen3.6 35b models at around 12tok/s. The prefill part is around 20tok/s when it works but can be hit and miss an