Nonparametric Learning of Two-Layer ReLU Residual Units
arXiv:2008.07648
Abstract
We describe an algorithm that learns two-layer residual units using rectified linear unit (ReLU) activation: suppose the input is from a distribution with support space and the ground-truth generative model is a residual unit of this type, given by , where ground-truth network parameters represent a full-rank matrix with nonnegative entries and is full-rank with and for , . We design layer-wise objectives as functionals whose analytic minimizers express the exact ground-truth network in terms of its parameters and nonlinearities. Following this objective landscape, learning residual units from finite samples can be formulated using convex optimization of a nonparametric function: for each layer, we first formulate the corresponding empirical risk minimization (ERM) as a positive semi-definite quadratic program (QP), then we show the solution space of the QP can be equivalently determined by a set of linear inequalities, which can then be efficiently solved by linear programming (LP). We further prove the strong statistical consistency of our algorithm, and demonstrate its robustness and sample efficiency through experimental results on synthetic data and a set of benchmark regression datasets.
Published in Transactions on Machine Learning Research (11/2022), slightly typographically revised
References in corpus (13)
- Sequence to Sequence Learning with Neural Networks
- Provable defenses against adversarial examples via the convex outer adversarial polytope
- Convergence Analysis of Two-layer Neural Networks with ReLU Activation
- High Dimensional Statistical Inference and Random Matrices
- Recovery Guarantees for One-hidden-layer Neural Networks
- Learning One-hidden-layer Neural Networks with Landscape Design
- What Can ResNet Learn Efficiently, Going Beyond Kernels?
- On the Computational Efficiency of Training Neural Networks
- An Analytical Formula of Population Gradient for two-layered ReLU network and its Applications in Convergence and Critical Point Analysis
- Efficient Learning of Generalized Linear and Single Index Models with Isotonic Regression
- Connecting Weighted Automata and Recurrent Neural Networks through Spectral Learning
- Learning Distributions Generated by One-Layer ReLU Networks
- Improved Learning of One-hidden-layer Convolutional Neural Networks with Overlaps