1 paper
Liu Ziyin, Yizhou Xu, Tomaso Poggio +1
Neural networks trained by gradient descent on a smooth cost function can nevertheless learn in steps: the cost holds on long plateaus and then drops abruptly. Meanwhile, training…