Skip to content
r/LocalLLaMA · Communities

Has anyone actually made 64k feel like 300k+ with recursive local agents?

I'm running Qwen 3.8 27B locally on a single GPU. I can push the context to 131k, but I'd rather run it faster at 64k if the agent can manage context properly. What I have in mind is pretty simple: one model stays loaded the whole time main agent gets 64k when something is too big, it spawns a fresh child with only the