Showing cs.LGShow all
3 papers · 1 filter
cs.LG2025
What Happens During the Loss Plateau? Understanding Abrupt Learning in Transformers
Pulkit Gopalani, Wei Hu
Training Transformers on algorithmic tasks frequently demonstrates an intriguing abrupt learning phenomenon: an extended performance plateau followed by a sudden, sharp improvement…
cs.LG2024
Abrupt Learning in Transformers: A Case Study on Matrix Completion
Pulkit Gopalani, Ekdeep Singh Lubana, Wei Hu
Recent analysis on the training dynamics of Transformers has unveiled an interesting characteristic: the training loss plateaus for a significant number of training steps, and then…
cs.LG2023
Global Convergence of SGD For Logistic Loss on Two Layer Neural Nets
Pulkit Gopalani, Samyak Jha, Anirbit Mukherjee
In this note, we demonstrate a first-of-its-kind provable convergence of SGD to the global minima of appropriately regularized logistic empirical risk of depth nets -- for arbi…