GLM-5.2-Int4-Int8 on 8× GB10: ~1,200 t/s prefill, 33–54 t/s avg decode
GLM-5.2-Int4-Int8 on 8× GB10: ~1,200 t/s prefill, 33–54 t/s avg decode (generic - coding/structured) and memory remaining to run also a Mimo 2.5 in parallel for…
GLM-5.2-Int4-Int8 on 8× GB10: ~1,200 t/s prefill, 33–54 t/s avg decode (generic - coding/structured) and memory remaining to run also a Mimo 2.5 in parallel for…
We have just gotten out of Ornith 2 weeks of spamming and now we start " 1 bit almost fp16 quality" BS. This is just pathetic.…
The current implementation of multi token prediction (MTP) in llama cpp could trigger BF16 compute selection on GPUs that don't support BF16, causing cuBLAS crashes on…
submitted by /u/Immediate_Sky_6566 [link] [comments]
I finally had a chance to play with AMD's new Ryzen AI Halo box, here I show my configuration that can get you 10-15% performance improvement…
I can how run 60 tk/s with two slot now, quality seems good, tool call is very stable. I haven't done any coding yet. 2 slot…
I've been experimenting with interpretability on Gemma-4-31B and ended up with something cool I think you guys might like: a variant that challenges a request's premise…
submitted by /u/yogthos [link] [comments]
Over past ~10 months I've been iterating on my memory system so I can make a proper assistant, like Rick's garage from Rick and Morty. I…
Very impressive release by the PrismML team. 1-bit quantization shrinks it from 54GB to just 3.8GB (-93%), while retaining 90% of its intelligence. - Collection on…