Running GLM 5.2 on 4xGB10 with a 100G Switch, 330k ctx, ~25 t/s tg, ~650 t/s pp
TP4+DCP2 for a ~360k kV pool. Prefill increases to 900-1000 t/s with longer prompts. You can also run DCP4 for 660k, but prefill gets shaved to…
TP4+DCP2 for a ~360k kV pool. Prefill increases to 900-1000 t/s with longer prompts. You can also run DCP4 for 660k, but prefill gets shaved to…
I'm glad to announce that we have completely revamped the site, so that you can now search the over 2000 viewable models instantly! You can also…
submitted by /u/panchovix [link] [comments]
Been using this for real software development for a commercial app. i.e. Not a single file HTML app. I mean a large scale 100k+ loc project…
It's perhaps the model AI model in the world right know, and I just wanted to understand, if you are serving a billion plus requests a…
Recently bought into the local LLM hype by buying a 32gb vram gpu and holy shit, gemma 4 31b at 5bits blows the standard ChatGPT model…
okay so you might be following me or not.. but I have been working in AI since last 10+ years and our first product in AI…
Hey all, I've been working with DSV4 on my Frankenstein rig (5-GPU box: 2x3090 + 5060 Ti + 2x4060 Ti, 96 GiB VRAM, 125 GiB DDR4)…
submitted by /u/beneath_steel_sky [link] [comments]
This PR adds Q2_0 support for CPU. Main motivation is to support Ternary Bonsai models (1.7B, 4B, 8B) and upcoming models. This PR is CPU only…