Skip to content
r/LocalLLaMA · Communities

CachyLLama’s: llama.cpp fork with persistent KV cache that makes long local-agent sessions much less painful

I’m not affiliated with this project, but I’ve been running it recently and I’m surprised it hasn’t received more attention here: https://github.com/fewtarius/CachyLLama CachyLLama is a fork of llama.cpp focused on a problem that matters a lot on slower hardware: repeated prompt processing. Not only does it have a new