1 paper
Tao Shi, Liangming Chen, Long Jin +1
In the training of neural networks, adaptive moment estimation (Adam) typically converges fast but exhibits suboptimal generalization performance. A widely accepted explanation for…