Robust Large Margin Deep Neural Networks
arXiv:1605.08254 · doi:10.1109/TSP.2017.2708039
Abstract
The generalization error of deep neural networks via their classification margin is studied in this work. Our approach is based on the Jacobian matrix of a deep neural network and can be applied to networks with arbitrary non-linearities and pooling layers, and to networks with different architectures such as feed forward networks and residual networks. Our analysis leads to the conclusion that a bounded spectral norm of the network's Jacobian matrix in the neighbourhood of the training samples is crucial for a deep neural network of arbitrary depth and width to generalize well. This is a significant improvement over the current bounds in the literature, which imply that the generalization error grows with either the width or the depth of the network. Moreover, it shows that the recently proposed batch normalization and weight normalization re-parametrizations enjoy good generalization properties, and leads to a novel network regularizer based on the network's Jacobian matrix. The analysis is supported with experimental results on the MNIST, CIFAR-10, LaRED and ImageNet datasets.
accepted to IEEE Transactions on Signal Processing
References in corpus (6)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Striving for Simplicity: The All Convolutional Net
- The Loss Surfaces of Multilayer Networks
- Path-SGD: Path-Normalized Optimization in Deep Neural Networks
- Norm-Based Capacity Control in Neural Networks
- Generalization Error of Invariant Classifiers
Cited by in corpus (100)
- Towards Deep Learning Models Resistant to Adversarial Attacks
- Wild Patterns: Ten Years After the Rise of Adversarial Machine Learning
- Sensitivity and Generalization in Neural Networks: an Empirical Study
- Generalization in Deep Learning
- Invertible Residual Networks
- Predicting the Generalization Gap in Deep Networks with Margin Distributions
- Simple and Principled Uncertainty Estimation with Deterministic Deep Learning via Distance Awareness
- Gradient Descent Maximizes the Margin of Homogeneous Neural Networks
- Why Do Adversarial Attacks Transfer? Explaining Transferability of Evasion and Poisoning Attacks
- MMA Training: Direct Input Space Margin Maximization through Adversarial Training
- Mathematics of Deep Learning
- Sorting out Lipschitz function approximation
- Controlling Neural Level Sets
- Procedural Noise Adversarial Examples for Black-Box Attacks on Deep Convolutional Networks
- Robust Learning with Jacobian Regularization
- On Robustness of Neural Ordinary Differential Equations
- Explicit Regularisation in Gaussian Noise Injections
- Generalization Error of Invariant Classifiers
- The Lipschitz Constant of Self-Attention
- Opportunities and Challenges in Deep Learning Adversarial Robustness: A Survey
- Improved Sample Complexities for Deep Networks and Robust Classification via an All-Layer Margin
- The Implicit and Explicit Regularization Effects of Dropout
- Lipschitz regularized Deep Neural Networks generalize and are adversarially robust
- Boundary thickness and robustness in learning models
- Heteroskedastic and Imbalanced Deep Learning with Adaptive Regularization
- Training robust and generalizable quantum models
- Understanding and Mitigating Exploding Inverses in Invertible Neural Networks
- A case for new neural network smoothness constraints
- Deep neural networks for choice analysis: Enhancing behavioral regularity with gradient regularization
- Robust learning with implicit residual networks
- Quantifying the generalization error in deep learning in terms of data distribution and neural network smoothness
- Generalization Error Bounds with Probabilistic Guarantee for SGD in Nonconvex Optimization
- Understanding Open-Set Recognition by Jacobian Norm and Inter-Class Separation
- Adversarial Margin Maximization Networks
- Theory IIIb: Generalization in Deep Networks
- Can We Gain More from Orthogonality Regularizations in Training Deep CNNs?
- Large Margin Deep Networks for Classification
- Discrete-Valued Neural Communication
- DNN or k-NN: That is the Generalize vs. Memorize Question
- Layer-wise Characterization of Latent Information Leakage in Federated Learning
- Self-supervised Neural Architecture Search
- Towards Robust Deep Neural Networks
- Jacobian Regularization for Mitigating Universal Adversarial Perturbations
- Towards Task and Architecture-Independent Generalization Gap Predictors
- Improving Transformation Invariance in Contrastive Representation Learning
- Orthogonal Deep Neural Networks
- Statistical Performance of Radio Interferometric Calibration
- PAC-Bayesian Margin Bounds for Convolutional Neural Networks
- Noisy Recurrent Neural Networks
- A Closer Look at Double Backpropagation
- Theory III: Dynamics and Generalization in Deep Networks
- Lipschitz Bounds and Provably Robust Training by Laplacian Smoothing
- Certifying Incremental Quadratic Constraints for Neural Networks via Convex Optimization
- Deep Learning for Inverse Problems: Bounds and Regularizers
- A Differential Topological View of Challenges in Learning with Feedforward Neural Networks
- Edge of chaos as a guiding principle for modern neural network training
- On the relationship between class selectivity, dimensionality, and robustness
- Information-Theoretic Local Minima Characterization and Regularization
- Deep Convolutional Framelet Denosing for Low-Dose CT via Wavelet Residual Network
- Towards Certifying L-infinity Robustness using Neural Networks with L-inf-dist Neurons
- A Simple yet Effective Baseline for Robust Deep Learning with Noisy Labels
- UnitedQA: A Hybrid Approach for Open Domain Question Answering
- Bridging the Gap Between Adversarial Robustness and Optimization Bias
- What training reveals about neural network complexity
- RoMA: Robust Model Adaptation for Offline Model-based Optimization
- An ETF view of Dropout regularization
- An Empirical Study on the Relation between Network Interpretability and Adversarial Robustness
- Towards a Theoretical Understanding of the Robustness of Variational Autoencoders
- Lipschitz Constrained GANs via Boundedness and Continuity
- Pre-interpolation loss behaviour in neural networks
- Neuron with Steady Response Leads to Better Generalization
- Margin-Based Regularization and Selective Sampling in Deep Neural Networks
- Ranking Neural Checkpoints
- Adaptive Prototypical Networks with Label Words and Joint Representation Learning for Few-Shot Relation Classification
- FeTa: A DCA Pruning Algorithm with Generalization Error Guarantees
- Revisiting hard thresholding for DNN pruning
- Posterior Differential Regularization with f-divergence for Improving Model Robustness
- Semantically Robust Unpaired Image Translation for Data with Unmatched Semantics Statistics
- The role of invariance in spectral complexity-based generalization bounds
- Generalisation in fully-connected neural networks for time series forecasting
- Gentle Local Robustness implies Generalization
- The Missing Margin: How Sample Corruption Affects Distance to the Boundary in ANNs
- Preprint: Norm Loss: An efficient yet effective regularization method for deep neural networks
- Analytical bounds on the local Lipschitz constants of affine-ReLU functions
- Dynamical System Inspired Adaptive Time Stepping Controller for Residual Network Families
- Noisy Feature Mixup
- Model-Aware Regularization For Learning Approaches To Inverse Problems
- On Connections between Regularizations for Improving DNN Robustness
- Topologically Densified Distributions
- Is the Meta-Learning Idea Able to Improve the Generalization of Deep Neural Networks on the Standard Supervised Learning?
- Recent Advances in Large Margin Learning
- Likelihood Landscapes: A Unifying Principle Behind Many Adversarial Defenses
- MadNet: Using a MAD Optimization for Defending Against Adversarial Attacks
- Linking average- and worst-case perturbation robustness via class selectivity and dimensionality
- GIM: Gaussian Isolation Machines
- Analytic expressions for the output evolution of a deep neural network
- Effect of Various Regularizers on Model Complexities of Neural Networks in Presence of Input Noise
- Bridged Adversarial Training
- Training Efficiency and Robustness in Deep Learning
- On Symmetry and Initialization for Neural Networks