1 paper · 1 filter
Yuhang Cai, Jingfeng Wu, Song Mei +2
The typical training of neural networks using large stepsize gradient descent (GD) under the logistic loss often involves two distinct phases, where the empirical risk oscillates i…