Skip to content
llama.cpp releases · Infrastructure

b9857

hexagon: flash attention rework (optimizations, accuracy improvements, etc) (#25085) hex-mm: fold mm quant tasks into the main matmul threads hex-mm: minor formatting fixes hex-mm: cleanup is_quant checks in dma dispatch hex-mm: fix dst-spad alignment hex-mm: move fp kernels in the hvx-mm-kernels header hex-mm: fuse wi