arXiv stat.ML
· Papers
Transformers with RL or SFT Provably Learn Sparse Boolean Functions, But Differently
arXiv:2511.17852v3 Announce Type: replace-cross Abstract: Transformers can acquire Chain-of-Thought (CoT) capabilities to solve reasoning tasks via fine-tuning. Reinforcement learning (RL) and supervised fine-tuning (SFT) are two primary approaches to this end. In this work, we examine RL with verifiable process reward