117 citations · 171 across the 6 of their papers we have counts for
3 papers · 1 filter
Wide flat minima and optimal generalization in classifying high-dimensional Gaussian mixtures
Carlo Baldassi, Enrico M. Malatesta, Matteo Negri +1
We analyze the connection between minimizers with good generalizing properties and high local entropy regions of a threshold-linear classifier in Gaussian mixtures with the mean sq…
Shaping the learning landscape in neural networks around wide flat minima
Carlo Baldassi, Fabrizio Pittorino, Riccardo Zecchina
Learning in Deep Neural Networks (DNN) takes place by minimizing a non-convex high-dimensional loss function, typically by a stochastic gradient descent (SGD) strategy. The learnin…
Parle: parallelizing stochastic gradient descent
Pratik Chaudhari, Carlo Baldassi, Riccardo Zecchina +3
We propose a new algorithm called Parle for parallel training of deep networks that converges 2-4x faster than a data-parallel implementation of SGD, while achieving significantly…