3 papers
cs.LG2025
What Happens During the Loss Plateau? Understanding Abrupt Learning in Transformers
Pulkit Gopalani, Wei Hu
Training Transformers on algorithmic tasks frequently demonstrates an intriguing abrupt learning phenomenon: an extended performance plateau followed by a sudden, sharp improvement…
cs.LG2024
Global Convergence of SGD On Two Layer Neural Nets
Pulkit Gopalani, Anirbit Mukherjee
In this note, we consider appropriately regularized empirical risk of depth nets with any number of gates and show bounds on how the empirical loss evolves for SGD ite…
cs.LG2024
Towards Size-Independent Generalization Bounds for Deep Operator Nets
Pulkit Gopalani, Sayar Karmakar, Dibyakanti Kumar +1
In recent times machine learning methods have made significant advances in becoming a useful tool for analyzing physical systems. A particularly active area in this theme has been…