4 papers · 1 filter
What Happens During the Loss Plateau? Understanding Abrupt Learning in Transformers
Pulkit Gopalani, Wei Hu
Training Transformers on algorithmic tasks frequently demonstrates an intriguing abrupt learning phenomenon: an extended performance plateau followed by a sudden, sharp improvement…
Global Convergence of SGD On Two Layer Neural Nets
Pulkit Gopalani, Anirbit Mukherjee
In this note, we consider appropriately regularized empirical risk of depth nets with any number of gates and show bounds on how the empirical loss evolves for SGD ite…
Towards Size-Independent Generalization Bounds for Deep Operator Nets
Pulkit Gopalani, Sayar Karmakar, Dibyakanti Kumar +1
In recent times machine learning methods have made significant advances in becoming a useful tool for analyzing physical systems. A particularly active area in this theme has been…
Abrupt Learning in Transformers: A Case Study on Matrix Completion
Pulkit Gopalani, Ekdeep Singh Lubana, Wei Hu
Recent analysis on the training dynamics of Transformers has unveiled an interesting characteristic: the training loss plateaus for a significant number of training steps, and then…