r/LocalLLaMA
· Communities
First Kimi K3 results on home lab ~ 4t/s
I've got better results than expected for 768gb DDR5 and 2x5090. Using fork https://github.com/pwilkin/llama.cpp/tree/kimi-k3-text and https://huggingface.co/GrEarl/Kimi-K3-GGUF Q2_K quant. Prefill speed for big prompt is 50-70 tps. The most fun thing that decoding tps growing over time. Maybe some kind of warmup or sw