Competence Gate: gating tool-use on a small model’s internal confidence signal instead of its verbalised one — Qwen3.5-4B, open weights [P]
I made a 10MB LoRA adapter for Qwen3.5-4B plus a small orchestration layer. It decides, per query, whether to answer directly, search the web, or retrieve…