Skip to content
r/LocalLLaMA · Communities

native 64k+ Gguf model for llama.cpp

I'm doing something a bit silly and trying to get hermes agent fully local on minimal resources. I've got a lenovop520 64gb quad channel ram and an rtx 3060. I'm using qwen3.5 35b with cpu mode so it only takes about 4g of my vram and still gets 30 t/s. I'm *trying* to configure a secondary model for LCM context compac