Revealing the Structure of Deep Neural Networks via Convex Duality
arXiv:2002.09773
Abstract
We study regularized deep neural networks (DNNs) and introduce a convex analytic framework to characterize the structure of the hidden layers. We show that a set of optimal hidden layer weights for a norm regularized DNN training problem can be explicitly found as the extreme points of a convex set. For the special case of deep linear networks, we prove that each optimal weight matrix aligns with the previous layers via duality. More importantly, we apply the same characterization to deep ReLU networks with whitened data and prove the same weight alignment holds. As a corollary, we also prove that norm regularized deep ReLU networks yield spline interpolation for one-dimensional datasets which was previously known only for two-layer networks. Furthermore, we provide closed-form solutions for the optimal layer weights when data is rank-one or whitened. The same analysis also applies to architectures with batch normalization even for arbitrary data. Therefore, we obtain a complete explanation for a recent empirical observation termed Neural Collapse where class means collapse to the vertices of a simplex equiangular tight frame.
Accepted to ICML 2021
References in corpus (5)
- Prevalence of Neural Collapse during the terminal phase of deep learning training
- The Role of Neural Network Activation Functions
- Width Provably Matters in Optimization for Deep Linear Neural Networks
- Convex Geometry and Duality of Over-parameterized Neural Networks
- Demystifying Batch Normalization in ReLU Networks: Equivalent Convex Optimization Models and Implicit Regularization
Cited by in corpus (11)
- Banach Space Representer Theorems for Neural Networks and Ridge Splines
- Neural Collapse Under MSE Loss: Proximity to and Dynamics on the Central Path
- Demystifying Batch Normalization in ReLU Networks: Equivalent Convex Optimization Models and Implicit Regularization
- Convex Geometry and Duality of Over-parameterized Neural Networks
- Neural Networks are Convex Regularizers: Exact Polynomial-time Convex Optimization Formulations for Two-layer Networks
- Convex Regularization Behind Neural Reconstruction
- Neural Spectrahedra and Semidefinite Lifts: Global Convex Optimization of Polynomial Activation Neural Networks in Fully Polynomial-Time
- Vector-output ReLU Neural Network Problems are Copositive Programs: Convex Analysis of Two Layer Networks and Polynomial-time Algorithms
- Hidden Convexity of Wasserstein GANs: Interpretable Generative Models with Closed-Form Solutions
- Training Quantized Neural Networks to Global Optimality via Semidefinite Programming
- Implicit Convex Regularizers of CNN Architectures: Convex Optimization of Two- and Three-Layer Networks in Polynomial Time