Skip to content
r/LocalLLaMA · Communities

60-82% accuracy swing on 4B model classification task: the only variable was harness design

I ran a pre-registered ablation on a classification task (Kubernetes issue → SIG triage) using a 4B model on a 6GB laptop GPU. Same frozen weights, same 250-issue gold corpus, same scorer across every run. The variable under test was harness design: rule placement, evidence order, turn structure, what survives between