Skip to content
HF Daily Papers · Papers

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization

Rubric-based reinforcement learning enriches language model training by evaluating model outputs against explicit criteria. Yet in GRPO-style pipelines, these structured judgments are reduced to a scalar response-level reward and converted into a response-level advantage, which is broadcast uniforml