8 citations · 24 across the 7 of their papers we have counts for
1 paper · 1 filter
Jingzhao Zhang, Tianxing He, Suvrit Sra +1
We provide a theoretical explanation for the effectiveness of gradient clipping in training deep neural networks. The key ingredient is a new smoothness condition derived from prac…