arXiv cs.AI
· Papers
Disentangling Feature Structure: A Mathematically Provable Two-Stage Training Dynamics in Transformers
arXiv:2502.20681v3 Announce Type: replace-cross Abstract: Transformers may exhibit two-stage training dynamics during the real-world training process. For instance, when training GPT-2 on the Counterfact dataset, the answers progress from syntactically incorrect to syntactically correct to semantically correct. However