RT vLLM: 🚀 @deepseek_ai's DSpark speculative decoding now runs natively in vLLM! What it is: a semi-autoregressive drafter that proposes several to…
RT vLLM🚀 @deepseek_ai's DSpark speculative decoding now runs natively in vLLM!What it is: a semi-autoregressive drafter that proposes several tokens in parallel with non-causal sliding-window attention,…