llama.cpp releases
· Infrastructure
b10447
server: re-design yield_to_queue thread model (#27133) run common_speculative_process in worker swap worker main thread design Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu ar