Skip to content
r/LocalLLaMA · Communities

llama.cpp slower on P-Cores than on E-Cores with MoE Model and GPU+CPU offloading?

I am currently experimenting with my setup: RTX 5090 + Intel 270K Plus CPU (8 Performance Cores + 16 Efficiency Cores) + 128 GB DDR5-6000 RAM, Ubuntu 26.04. I wanted to test the performance of Qwen 3.5 122b a10b with CPU offloading. Some mentioned that pinning llama.cpp to CPU performance cores could improve performanc