3 papers
cs.CL2025
From Dense to Dynamic: Token-Difficulty Driven MoEfication of Pre-Trained LLMs
Kumari Nishu, Sachin Mehta, Samira Abnar +6
Training large language models (LLMs) for different inference constraints is computationally expensive, limiting control over efficiency-accuracy trade-offs. Moreover, once trained…
cs.CL2024
Scaling Smart: Accelerating Large Language Model Pre-training with Small Model Initialization
Mohammad Samragh, Iman Mirzadeh, Keivan Alizadeh Vahid +5
The pre-training phase of language models often begins with randomly initialized parameters. With the current trends in scaling models, training their large number of parameters ca…
eess.AS2024
SLiCK: Exploiting Subsequences for Length-Constrained Keyword Spotting
Kumari Nishu, Minsik Cho, Devang Naik
User-defined keyword spotting on a resource-constrained edge device is challenging. However, keywords are often bounded by a maximum keyword length, which has been largely under-le…