36 citations · 48 across the 2 of their papers we have counts for
2 papers
cs.LG2022★ 36 cited
Understanding AdamW through Proximal Methods and Scale-Freeness
Zhenxun Zhuang, Mingrui Liu, Ashok Cutkosky +1
Adam has been widely adopted for training deep neural networks due to less hyperparameter tuning and remarkable performance. To improve generalization, Adam is typically used in ta…
cs.LG2020★ 12 cited
Adam: A Stochastic Method with Adaptive Variance Reduction
Mingrui Liu, Wei Zhang, Francesco Orabona +1
Adam is a widely used stochastic optimization method for deep learning applications. While practitioners prefer Adam because it requires less parameter tuning, its use is problemat…