3 papers
cs.LG2026
pscaling small models: Principled warm starts and hyperparameter transfer
Yuxin Ma, Nan Chen, Mateo DÃaz +3
Modern large-scale neural networks are often trained and released in multiple sizes to accommodate diverse inference budgets. To improve efficiency, recent work has explored model…
cs.LG2026
On the Stability of the Jacobian Matrix in Deep Neural Networks
Benjamin Dadoun, Soufiane Hayou, Hanan Salam +2
Deep neural networks are known to suffer from exploding or vanishing gradients as depth increases, a phenomenon closely tied to the spectral behavior of the input-output Jacobian.…
cs.LG2025
Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size
Soufiane Hayou, Liyuan Liu
Pretraining large language models is a costly process. To make this process more efficient, several methods have been proposed to optimize model architecture/parametrization and ha…