Skip to content
r/LocalLLaMA · Communities

Has anyone actually benchmarked where the "big-model orchestrator + local-model worker" split breaks down?

I keep seeing the "use a big model via API as the architect, run local small/mid models as workers" pattern recommended for people with modest local hardware. I've been running it myself (orchestrator on a hosted model, local Qwen-class 27B workers doing scans/refactors/test runs), and it works - but I have a nagging f