Skip to content
r/LocalLLaMA · Communities

VRAM disk cache of MoE makes 340 pp/s 9.6 tg/s for Kimi 2.7 on a single dgx spark

this strategy effectively uses vram as cache over disk to keep MoE experts on cuda compute path in llama.cpp. numbers first. detailed explanation down below. Numbers on dgx spark Kimi-K2.7-Code.i1-IQ_S.gguf 204GB 1T.A32B https://huggingface.co/mradermacher/Kimi-K2.7-Code-i1-GGUF run method pp512 tg128 note A cuda, -ot