Skip to content
arXiv cs.CL · Papers

Complexity-Guided Component-wise Initialization for Language Model Pretraining

arXiv:2607.09204v1 Announce Type: new Abstract: Pretrained language models often exhibit structured weight spectra, suggesting that training may repeatedly produce similar layerwise and component-wise organization. We ask whether these recurring spectral patterns can be reused as an initialization signal for GPT-2-styl