Skip to content
r/LocalLLaMA · Communities

Quantizing Kimi K3 (2.8T A50B) to GGUF ourselves – Q3_K_S works, 1.1 TB on disk

we're experimenting with our own dynamic GGUF quants of kimi k3, made from the original weights with our llama.cpp fork. Q3_K_S is done and works 1114.76 GiB on disk. Q1 and Q2 are in progress, results on those tomorrow rented box hardware: - AMD EPYC 9554P, 64 cores - 1.5 TB of DDR5 - NVMe in raid0 to store the weight