r/LocalLLaMA
· Communities
Has anyone actually benchmarked where the "big-model orchestrator + local-model worker" split breaks down?
I keep seeing the "use a big model via API as the architect, run local small/mid models as workers" pattern recommended for people with modest local hardware. I've been running it myself (orchestrator on a hosted model, local Qwen-class 27B workers doing scans/refactors/test runs), and it works - but I have a nagging f