Skip to content
llama.cpp releases · Infrastructure

b10355

llama : support multi-output backend sampling (#25532) Enable backend sampling with token speculation Clamp the mask sum before converting it into the sampled index Add a numeric context parameter declaring the maximum outputs one sequence More fixes Don't reuse memory for output views. Match dist between CPU and GPU F