Skip to content
r/LocalLLaMA · Communities

How i got Bonsai-Ternary-27B to run at 120k context <10gb vram.

Hey guys, so its quite a read, its long. If you are bored just read the TLDR and see if its worth it for you: TLDR; Asked Deepseek to port KVarN paper into PrismML's Bonsai runtime. Ended up +68% faster (73 vs 43 tok/s) and 3.3GB less VRAM at 120K context. So PrismML just dropped that Bonsai-27B 1-bit 3.8GB and a sligh