Skip to content
r/LocalLLaMA · Communities

Bonsai 27B runs locally on an iPhone – a 27B model in 3.9GB

PrismML built Bonsai on top of Qwen3.6-27B by quantizing the weights down to 1-bit. That takes it from ~54GB to 3.9GB, small enough to fit and run on a phone, while keeping ~90% of the benchmark scores It's true binary quantization ("binary g128") - every weight is a single sign bit and each group of 128 shares one FP1