3 papers
cs.LG2022
Restricted Strong Convexity of Deep Learning Models with Smooth Activations
Arindam Banerjee, Pedro Cisneros-Velarde, Libin Zhu +1
We consider the problem of optimization of deep learning models with smooth activation functions. While there exist influential results on the problem from the ``near initializatio…
cs.LG2022
Transition to Linearity of Wide Neural Networks is an Emerging Property of Assembling Weak Models
Chaoyue Liu, Libin Zhu, Mikhail Belkin
Wide neural networks with linear output layer have been shown to be near-linear, and to have near-constant neural tangent kernel (NTK), in a region containing the optimization path…
cs.LG2020
On the linearity of large non-linear models: when and why the tangent kernel is constant
Chaoyue Liu, Libin Zhu, Mikhail Belkin
The goal of this work is to shed light on the remarkable phenomenon of transition to linearity of certain neural networks as their width approaches infinity. We show that the trans…