Skip to content
arXiv cs.CL · Papers

Behavior Leverage Imbalance in Multi-Teacher On-Policy Distillation

arXiv:2607.07050v1 Announce Type: new Abstract: Agentic language models must learn when to call tools, when to consume tool responses, and when to answer directly. This makes multi-teacher on-policy distillation a natural training strategy: one teacher can specialize in tool calls, another in direct responses, and the