Skip to content
r/LocalLLaMA · Communities

SOTA Apple Silicon Inference (August 15, 2026)

This is handwritten. Enjoy. TL;DR I've spent the last 2 weeks full-time looking into the state of inference optimization on Apple Silicon, and honestly, the software stack is a mess. There is no framework that has all the inference optimizations that are available on CUDA/NVIDIA for the latest Qwen models: prefix cachi