r/LocalLLaMA
· Communities
[Release] WinterMix — 3 Bit WinterMix of Qwen3.5-122B-A10B in native MLX: a 59 GiB build with best-in-class Long Context coherence
TL;DR: I spent another 8 days following my last post making major improvements to the WinterMix method for MLX models. At 20k+ context this 59 GiB build posts a better perplexity than even UnSloth's Q3_K_XL GGUF* thanks to the new annealing process on its reasoning traces** (new at the wMix38 tier, not yet applied to p