Skip to content
arXiv stat.ML · Papers

Transformers with RL or SFT Provably Learn Sparse Boolean Functions, But Differently

arXiv:2511.17852v3 Announce Type: replace-cross Abstract: Transformers can acquire Chain-of-Thought (CoT) capabilities to solve reasoning tasks via fine-tuning. Reinforcement learning (RL) and supervised fine-tuning (SFT) are two primary approaches to this end. In this work, we examine RL with verifiable process reward