4 papers
From Dense to Dynamic: Token-Difficulty Driven MoEfication of Pre-Trained LLMs
Kumari Nishu, Sachin Mehta, Samira Abnar +6
Training large language models (LLMs) for different inference constraints is computationally expensive, limiting control over efficiency-accuracy trade-offs. Moreover, once trained…
M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference
Nikhil Bhendawade, Mahyar Najibi, Devang Naik +1
Residual transformations enhance the representational depth and expressive power of large language models (LLMs). However, applying static residual transformations across all token…
Scaling Smart: Accelerating Large Language Model Pre-training with Small Model Initialization
Mohammad Samragh, Iman Mirzadeh, Keivan Alizadeh Vahid +5
The pre-training phase of language models often begins with randomly initialized parameters. With the current trends in scaling models, training their large number of parameters ca…
SLiCK: Exploiting Subsequences for Length-Constrained Keyword Spotting
Kumari Nishu, Minsik Cho, Devang Naik
User-defined keyword spotting on a resource-constrained edge device is challenging. However, keywords are often bounded by a maximum keyword length, which has been largely under-le…