arXiv cs.LG
· Papers
APPO: Agentic Procedural Policy Optimization
arXiv:2606.12384v2 Announce Type: replace Abstract: Recent advances in agentic Reinforcement Learning (RL) have substantially improved the multi-turn tool-use capabilities of large language model agents. However, most existing methods assign credit over coarse heuristic units, such as tool-call boundaries or fixed work