Skip to content
r/LocalLLaMA · Communities

Meta’s Muse Glimmer 30B now runs up to ~3.3x faster on Mac with mlx-dspark

Been tinkering with speculative decoding on Apple Silicon for a while, and this week I got Meta's new Muse Glimmer 30B working in my project mlx-dspark. On my M4 Pro, the 8-bit model goes from 8.2 tok/s to 18-26 tok/s depending on content. Math is the best case at 3.27x, code 2.5x, chat 2.22x. Output is byte-identical