Skip to content
llama.cpp releases · Infrastructure

b10032

cuda : CUDA GGML_OP_LIGHTNING_INDEXER implementation (generic vector kernel + wmma kernel) (#25545) cuda : CUDA GGML_OP_LIGHTNING_INDEXER implementation (generic vector kernel + wmma kernel) chore : remove indentation of #pragma unroll cuda : remove unnecessary kernel template declarations cuda : add WARPS_PER_BLOCK an