Skip to content
llama.cpp releases · Infrastructure

b9952

llama : make all KQ masks f16 if FA is used, remove zero attention bias, remove raw_k repeats in DeepSeek V4 (#25370) llama : make all KQ masks (except the lightning indexer one) f16 if FA is used and remove zero attention bias in DeepSeek V4 llama : remove dead code that repeats unified raw_k cache for each stream in