arXiv cs.CL
· Papers
Complexity-Guided Component-wise Initialization for Language Model Pretraining
arXiv:2607.09204v1 Announce Type: new Abstract: Pretrained language models often exhibit structured weight spectra, suggesting that training may repeatedly produce similar layerwise and component-wise organization. We ask whether these recurring spectral patterns can be reused as an initialization signal for GPT-2-styl