Skip to content
r/LocalLLaMA · Communities

Need help tuning cache in llama-server

Hey I am running a few models on a strix halo box. Especially for the larger models (like Qwen 3.5 122B) they okayish performance wise if the cache is utilized properly but a full cache miss at 100k context causes roughly 10-20 minute of PP time - which is extremely annoying. I will first show what I have already confi