llama.cpp releases
· Infrastructure
b9910
server : fix draft model fit vs load inconsistency (#25056) fix: draft model fit vs load inconsistency refactor(server): unify draft/mtp parameter initialization, model, and context load moves speculative init to speculative.cpp changes server_context_impl model_dft and ctx_dft to use raw pointers fix: don't throttle p