r/LocalLLaMA
· Communities
I added MTP to local SoTA Agentic Coding Model Ornith 35B FP8 E4M3
Just wanted to share that I was looking for an optimal way to run Ornith 35B in FP8 with E4M3 and MTP with vLLM but there was no out-of-the-box model with MTP drafter support. So I grafted this new model! It's 18% faster than without MTP and the drafter acceptance rate is not bad (70% on avg). It should run on any RTX