1 paper
Mohammad Samragh, Iman Mirzadeh, Keivan Alizadeh Vahid +5
The pre-training phase of language models often begins with randomly initialized parameters. With the current trends in scaling models, training their large number of parameters ca…