llama.cpp releases
· Infrastructure
b10355
llama : support multi-output backend sampling (#25532) Enable backend sampling with token speculation Clamp the mask sum before converting it into the sampled index Add a numeric context parameter declaring the maximum outputs one sequence More fixes Don't reuse memory for output views. Match dist between CPU and GPU F