Skip to content
arXiv cs.CL · Papers

RL Post-Training Builds Compositional Reasoning Strategies

arXiv:2607.07646v1 Announce Type: cross Abstract: Does RL post-training merely amplify primitive skills already latent in a base model, or can it compose primitive skills into new higher-level strategies? We study this question in a fully observable rewrite-grammar environment where the pretraining distribution is know