Skip to content
llama.cpp releases · Infrastructure

b9893

opencl: general flash attention decode performance optimizations (#25366) opencl: vec flash-attention decode kernels for f16/q8_0/q4_0 KV opencl: improve non FA KQ mv kernels opencl: tweaks for multiquery FA opencl: some tweaks for FA q1 kernels opencl: FA with DK=DV=512 for gemma-4 opencl: various fixes opencl: cleanu