Skip to content
r/LocalLLaMA · Communities

llama.cpp: Hy3 PR + GGUFs

Early stages as the model was just released yesterday, but seems to be working already. Yay! Getting coherent output from the Q2_K, at about 10-11t/s on a 5090 + Zen 4 w/ 96GB DDR5. https://github.com/ggml-org/llama.cpp/pull/25395 https://huggingface.co/satgeze/Hy3-1M-GGUF Thanks to the PR author, satindergrewal! submi