Skip to content
r/LocalLLaMA · Communities

The harness matters more than the model. A 27B behind good critics changed my mind.

I saw someone test Qwen3.6-27B with a 3-critic harness. The harness included code review, test review and Playwright e2e. Each critic had context. The result was that the model is usable for coding work. This matches what I have come to believe from running agents in production. The harness around the model is more imp