Skip to content
r/LocalLLaMA · Communities

Qwen3.8-27B is now up to ~3× faster on Apple Silicon with mlx-dspark

mlx-dspark is an MLX port of DeepSeek's DSpark speculative-decoding drafters (the DeepSpec release), plus z-lab's DFlash, with one lossless verify loop. v0.10.0 adds Qwen3.8-27B via RadixArk's drafter, the first SpecForge/SGLang-packaged head it loads. Numbers (M4 Pro 48 GB, medians of 3, greedy, output ids identical t