Skip to content
arXiv cs.LG · Papers

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks

arXiv:2607.22724v1 Announce Type: new Abstract: Group-based policy optimization has been increasingly used to train large language model (LLM) agents from sparse outcome rewards by comparing trajectories or steps within a group. However, on difficult long-horizon tasks, this comparison can suffer from a sampling imbala