r/LocalLLaMA
· Communities
Thinking about grabbing 4x Ascend GX10s
Some in this sub have tested GLM5.2 on 4x DGX Sparks (or Ascend GX10) with 400-500 tok/s prompt processing and ~15 tok/s output at 128k context. Not blazing fast, but usable imo, especially with quantization. My thinking: If there's an open-source fable 5 sometime in december or next year, I would rather already have h