arXiv cs.CL
· Papers
DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment
arXiv:2607.07820v1 Announce Type: new Abstract: Training tool-use agents to improve from their own experience remains challenging, as supervised fine-tuning relies on fixed teacher-distilled trajectories, while sparse-reward reinforcement learning provides weak supervision for long-horizon interactions. We present Deep