Skip to content
arXiv cs.AI · Papers

In-the-Flow Agentic System Optimization for Effective Planning and Tool Use

arXiv:2510.05592v2 Announce Type: replace Abstract: Outcome-driven reinforcement learning has advanced reasoning in large language models (LLMs), but prevailing tool-augmented approaches train a single, monolithic policy that interleaves thoughts and tool calls under full context; this scales poorly with long horizons