1 paper
Tom Sander, Maxime Sylvestre, Alain Durmus
Training Deep Neural Networks (DNNs) with small batches using Stochastic Gradient Descent (SGD) yields superior test performance compared to larger batches. The specific noise stru…