llama.cpp releases
· Infrastructure
b10342
model : Granite-Switch Architecture (#25107) granite-switch: add llama.cpp backend (POC, CPU) New "granite-switch" architecture: a dense, all-attention Granite-4.1 model with N embedded LoRA adapters selected per-token by control tokens. gguf-py schema (arch, KV keys, stacked LoRA tensor names) + writer helpers convers