r/LocalLLaMA
· Communities
I got Nemotron Puzzle 75B running smoothly on a 64GB M2 Max
TL;DR: Added native nemotron_h_puzzle support to mlx-lm (PR #1535), then compared 4-bit vs 5-bit expert quantization (both with 6-bit dense layers, BF16 output head, group size 64) on a 64GB M2 Max. Results (same prompts, 5 seeds per task, temp 1.0 / top_p 0.95): 4-bit experts 5-bit experts Dense paths 6-bit 6-bit Outp