Skip to content
r/LocalLLaMA · Communities

cuda: extract Q1_0 elements via __byte_perm by dfriehs · Pull Request #25628 · ggml-org/llama.cpp

5% boost on tg for Bonsai models. This is the 1st PR mentioned on yesterday thread(Other Open PRs section) submitted by /u/pmttyji [link] [comments]