1.6k citations · 1.7k across the 9 of their papers we have counts for
1 paper · 1 filter
Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit +4
While stochastic gradient descent (SGD) is still the \emph{de facto} algorithm in deep learning, adaptive methods like Clipped SGD/Adam have been observed to outperform SGD across…