6 citations · 6 across the 2 of their papers we have counts for
2 papers
cs.LG2021★ 6 cited
A Loss Curvature Perspective on Training Instability in Deep Learning
Justin Gilmer, Behrooz Ghorbani, Ankush Garg +6
In this work, we study the evolution of the loss Hessian across many classification tasks in order to understand the effect the curvature of the loss has on the training dynamics.…
cs.CL2021
Beyond Distillation: Task-level Mixture-of-Experts for Efficient Inference
Sneha Kudugunta, Yanping Huang, Ankur Bapna +4
Sparse Mixture-of-Experts (MoE) has been a successful approach for scaling multilingual translation models to billions of parameters without a proportional increase in training com…