Skip to content
llama.cpp releases · Infrastructure

b9993

model: add Hy3 (hy_v3) support with MTP speculative decoding (#25395) model: add Hy3 (hy_v3) architecture support Adds Tencent Hunyuan 3 (HF architecture HYV3ForCausalLM, GGUF arch hy_v3): a MoE decoder stack with per-head Q/K RMSNorm, a sigmoid router with expert selection bias, an always-active ungated shared expert,