2 citations · 2 across the 2 of their papers we have counts for
2 papers
cs.CL2024
Scaling Smart: Accelerating Large Language Model Pre-training with Small Model Initialization
Mohammad Samragh, Iman Mirzadeh, Keivan Alizadeh Vahid +5
The pre-training phase of language models often begins with randomly initialized parameters. With the current trends in scaling models, training their large number of parameters ca…
cs.CL2024★ 2 cited
OpenELM: An Efficient Language Model Family with Open Training and Inference Framework
Sachin Mehta, Mohammad Hossein Sekhavat, Qingqing Cao +8
The reproducibility and transparency of large language models are crucial for advancing open research, ensuring the trustworthiness of results, and enabling investigations into dat…