r/LocalLLaMA
· Communities
Our 1-bit quant of Hy3 295B runs 2.2x faster than the cloud API with no quality loss
We quantized Tencent's Hy3 295B down to 1 bit and got a 92GB IQ1_M GGUF, small enough for one 4-GPU box. We ran it on 4x RTX 5090 against the same Hy3 over the cloud API. Both got the same one-shot task. Each model built a self-playing retro game in one HTML file. We ran three rounds with Flappy Bird, Arkanoid and Snak