Skip to content
r/LocalLLaMA · Communities

Anyone tried using the new (ish) Gemma diffusion model as a speculative model?

It seems that MTP is the gold standard for speed up but still suffers from having to choose between regressive and parallel drafters that come with trade offs. Wouldn't using a diffusion model be a good way to get a quick 256 token draft of high quality? submitted by /u/Demonicated [link] [comments]