The Local Elasticity of Neural Networks
arXiv:1910.06943
Abstract
This paper presents a phenomenon in neural networks that we refer to as \textit{local elasticity}. Roughly speaking, a classifier is said to be locally elastic if its prediction at a feature vector $\bx'$ is \textit{not} significantly perturbed, after the classifier is updated via stochastic gradient descent at a (labeled) feature vector $\bx$ that is \textit{dissimilar} to $\bx'$ in a certain sense. This phenomenon is shown to persist for neural networks with nonlinear activation functions through extensive simulations on real-life and synthetic datasets, whereas this is not observed in linear classifiers. In addition, we offer a geometric interpretation of local elasticity using the neural tangent kernel \citep{jacot2018neural}. Building on top of local elasticity, we obtain pairwise similarity measures between feature vectors, which can be used for clustering in conjunction with -means. The effectiveness of the clustering algorithm on the MNIST and CIFAR-10 datasets in turn corroborates the hypothesis of local elasticity of neural networks on real-life data. Finally, we discuss some implications of local elasticity to shed light on several intriguing aspects of deep neural networks.
To appear in ICLR 2020
References in corpus (8)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- The Loss Surfaces of Multilayer Networks
- Stochastic Gradient Descent Optimizes Over-parameterized Deep ReLU Networks
- Towards moderate overparameterization: global convergence guarantees for training shallow neural networks
- Globally Optimal Gradient Descent for a ConvNet with Gaussian Inputs
- Complexity of Linear Regions in Deep Networks
- Classification regions of deep neural networks
- Critical Points of Neural Networks: Analytical Forms and Landscape Properties