Skip to content
llama.cpp releases · Infrastructure

b10447

server: re-design yield_to_queue thread model (#27133) run common_speculative_process in worker swap worker main thread design Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu ar