Skip to content
llama.cpp releases · Infrastructure

b10414

metal : add TQ2_0 support (#26980) metal: add TQ2_0 support Add support for the GGML_TYPE_TQ2_0 (ternary, 2 bits per element) type in the Metal backend. Assisted-by: llama.cpp:DeepSeek-v4-Flash-0731 cont : optimize mul_mv kernel float ops over integer ops precalculate sums hoist coef out of the inner loop contiguous y