6 papers
A Scalable Measure of Loss Landscape Curvature for Analyzing the Training Dynamics of LLMs
Dayal Singh Kalra, Jean-Christophe Gagnon-Audet, Andrey Gromov +4
Understanding the curvature evolution of the loss landscape is fundamental to analyzing the training dynamics of neural networks. The most commonly studied measure, Hessian sharpne…
Teaching Pretrained Language Models to Think Deeper with Retrofitted Recurrence
Sean McLeish, Ang Li, John Kirchenbauer +7
Recent advances in depth-recurrent language models show that recurrence can decouple train-time compute and parameter count from test-time compute. In this work, we study how to co…
When Can You Get Away with Low Memory Adam?
Dayal Singh Kalra, John Kirchenbauer, Maissam Barkeshli +1
Adam is the go-to optimizer for training modern machine learning models, but it requires additional memory to maintain the moving averages of the gradients and their squares. While…
(How) Can Transformers Predict Pseudo-Random Numbers?
Tao Tao, Darshil Doshi, Dayal Singh Kalra +2
Transformers excel at discovering patterns in sequential data, yet their fundamental limitations and learning mechanisms remain crucial topics of investigation. In this paper, we s…
Initializing ReLU networks in an expressive subspace of weights
Dayal Singh, G J Sreejith
Using a mean-field theory of signal propagation, we analyze the evolution of correlations between two signals propagating forward through a deep ReLU network with correlated weight…
Automated Detection of Solar Radio Bursts using a Statistical Method
Dayal Singh, K. Sasikumar Raja, Prasad Subramanian +2
Radio bursts from the solar corona can provide clues to forecast space weather hazards. After recent technology advancements, regular monitoring of radio bursts has increased and l…