Skip to content
arXiv cs.AI · Papers

CAST: Game Solvers as Turn-Level Teachers for LLM Agents

arXiv:2607.25308v1 Announce Type: cross Abstract: Training large language models (LLMs) to act in long-horizon games is a promising step toward generalist decision-making, yet reinforcement learning with verifiable rewards (RLVR) relies on sparse final rewards that reveal little about which decisions determine success.