Skip to content
X · @huggingface · X / Twitter

RT Georgi Gerganov: llama.cpp recently added DFlash support to its speculative decoding arsenal. Along with MTP, Eagle3 and various ngram-based techni…

RT Georgi Gerganovllama.cpp recently added DFlash support to its speculative decoding arsenal. Along with MTP, Eagle3 and various ngram-based techniques, the local model performance takes another step up.Special thanks to NVIDIA team and Ruixiang Wang specifically for leading this effort!https://github.com/ggml-org/lla