Skip to content
llama.cpp releases · Infrastructure

b10121: ui: reduce per-token render cost when streaming (#26053)

performance harness - the empirical root Assisted-by: Claude Opus 4.8 210.36ms -> 2.67ms per streamed token Assisted-by: Claude Opus 4.8 11.58ms -> 0.62ms per streamed token Assisted-by: Claude Opus 4.8 22.02ms -> 3.33ms per streamed token Assisted-by: Claude Opus 4.8 3.07ms -> 1.36ms per streamed token at 40 messages