r/LocalLLaMA
· Communities
cuda: extract Q1_0 elements via __byte_perm by dfriehs · Pull Request #25628 · ggml-org/llama.cpp
5% boost on tg for Bonsai models. This is the 1st PR mentioned on yesterday thread(Other Open PRs section) submitted by /u/pmttyji [link] [comments]