Skip to content
r/LocalLLaMA · Communities

model: add Hy3 (hy_v3) support with MTP speculative decoding by satindergrewal · Pull Request #25395 · ggml-org/llama.cpp

Adds support for Tencent's Hy3 (hy_v3 / HYV3ForCausalLM, 299B MoE, 80 layers + 1 MTP layer), including its multi-token-prediction head as a draft-mtp speculative target. submitted by /u/pmttyji [link] [comments]