r/LocalLLaMA
· Communities
Ternary Bonsai 1.58-bit models – ggml: add Q2_0 quantization support (CPU) by khosravipasha · Pull Request #24448 · ggml-org/llama.cpp
This PR adds Q2_0 support for CPU. Main motivation is to support Ternary Bonsai models (1.7B, 4B, 8B) and upcoming models. This PR is CPU only (ARM NEON + generic scalar fallback). This completes the Q1_0, Q2_0, Q4_0, Q8_0 family. We have the x86, Metal, CUDA, and Vulkan backends ready to submit later. https://huggingf