arXiv cs.CL
· Papers
Agon: Competitive Cross-Model RL with Implicit Rival Grading of Reasoning
arXiv:2607.07690v1 Announce Type: cross Abstract: Reinforcement learning from verifiable rewards (e.g. GRPO) is the engine behind today's reasoning models, yet it grades only the final answer. On hard problems this trains models to write more rather than to think better, since the trace itself is never graded and no la