1 paper
Hongxu Chen, Ke Wei, Xiaoming Yuan +1
The empirical evidence indicates that stochastic optimization with heavy-tailed gradient noise is more appropriate to characterize the training of machine learning models than that…