Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
From Dense to Dynamic: Token-Difficulty Driven MoEfication of Pre-Trained LLMs
Kumari Nishu, Sachin Mehta, Samira Abnar +6
Training large language models (LLMs) for different inference constraints is computationally expensive, limiting control over efficiency-accuracy trade-offs. Moreover, once trained…
cs.CL2024
Scaling Smart: Accelerating Large Language Model Pre-training with Small Model Initialization
Mohammad Samragh, Iman Mirzadeh, Keivan Alizadeh Vahid +5
The pre-training phase of language models often begins with randomly initialized parameters. With the current trends in scaling models, training their large number of parameters ca…