Skip to content
r/LocalLLaMA · Communities

mistral.rs v0.9.0: up to 1.8x faster CPU decode than llama.cpp on x86 and ARM!

https://preview.redd.it/nuk5rxceptbh1.png?width=1448&format=png&auto=webp&s=300344dd4c6552379e8536b81ba288be3d6dca3f On Qwen3 4B Q4_K, mistral.rs decodes faster than llama.cpp at every context depth we measured, on x86 (Sapphire Rapids) and ARM (GB10). We optimized mistral.rs at granular levels to achieve general speed