most citedTheory of Deep Learning III: explaining the non-overfitting puzzle

49 citations · 93 across the 2 of their papers we have counts for

collaborators
Showing cs.LGShow all

5 papers · 1 filter

cs.LG2019

Theory III: Dynamics and Generalization in Deep Networks

Andrzej Banburski, Qianli Liao, Brando Miranda +4

The key to generalization is controlling the complexity of the network. However, there is no obvious control of complexity -- such as an explicit regularization term -- in the trai…

cs.LG2018

A Surprising Linear Relationship Predicts Test Performance in Deep Networks

Qianli Liao, Brando Miranda, Andrzej Banburski +2

Given two networks with the same training loss on a dataset, when would they have drastically different test losses and errors? Better understanding of this question of generalizat…

cs.LG2018

Theory IIIb: Generalization in Deep Networks

Tomaso Poggio, Qianli Liao, Brando Miranda +3

A main puzzle of deep neural networks (DNNs) revolves around the apparent absence of "overfitting", defined in this paper as follows: the expected error does not get worse when inc…

cs.LG201849 cited

Theory of Deep Learning III: explaining the non-overfitting puzzle

Tomaso Poggio, Kenji Kawaguchi, Qianli Liao +5

A main puzzle of deep networks revolves around the absence of overfitting despite large overparametrization and despite the large capacity demonstrated by zero training error on ra…

cs.LG201844 cited

Theory of Deep Learning IIb: Optimization Properties of SGD

Chiyuan Zhang, Qianli Liao, Alexander Rakhlin +3

In Theory IIb we characterize with a mix of theory and experiments the optimization of deep convolutional networks by Stochastic Gradient Descent. The main new result in this paper…