5 citations · 6 across the 8 of their papers we have counts for
8 papers · 1 filter
Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate
Dayal Singh Kalra, Maissam Barkeshli
Hyperparameter transfer allows extrapolating optimal optimization hyperparameters from small to large scales, making it critical for training large language models (LLMs). This is…
A Scalable Measure of Loss Landscape Curvature for Analyzing the Training Dynamics of LLMs
Dayal Singh Kalra, Jean-Christophe Gagnon-Audet, Andrey Gromov +4
Understanding the curvature evolution of the loss landscape is fundamental to analyzing the training dynamics of neural networks. The most commonly studied measure, Hessian sharpne…
When Can You Get Away with Low Memory Adam?
Dayal Singh Kalra, John Kirchenbauer, Maissam Barkeshli +1
Adam is the go-to optimizer for training modern machine learning models, but it requires additional memory to maintain the moving averages of the gradients and their squares. While…
(How) Can Transformers Predict Pseudo-Random Numbers?
Tao Tao, Darshil Doshi, Dayal Singh Kalra +2
Transformers excel at discovering patterns in sequential data, yet their fundamental limitations and learning mechanisms remain crucial topics of investigation. In this paper, we s…
Why Warmup the Learning Rate? Underlying Mechanisms and Improvements
Dayal Singh Kalra, Maissam Barkeshli
It is common in deep learning to warm up the learning rate , often by a linear schedule between and a predetermined target . In this paper…
Universal Sharpness Dynamics in Neural Network Training: Fixed Point Analysis, Edge of Stability, and Route to Chaos
Dayal Singh Kalra, Tianyu He, Maissam Barkeshli
In gradient descent dynamics of neural networks, the top eigenvalue of the loss Hessian (sharpness) displays a variety of robust phenomena throughout training. This includes early…