Skip to content
r/LocalLLaMA · Communities

Local models + big context = slow. How are you orchestrating "map-reduce" style agent workflows?

I tried running local models (qwen3.6*, ds4 flash, gemma4*, etc) on my mbp pro m5 with 128Gb of unified memory and concluded the bottleneck is context size. The moment a conversation gets long (16k is already the bottleneck), inference slows to a crawl. If you work with Hermes agent you know this context size is almost