Skip to content
r/LocalLLaMA · Communities

Weight-Aware Streaming Tensor Engine: run Kimi K3 using 29 GB of RAM at 0.50 tok/s

submitted by /u/galapag0 [link] [comments]