On the Number of Linear Regions of Deep Neural Networks
arXiv:1402.1869
Abstract
We study the complexity of functions computable by deep feedforward neural networks with piecewise linear activations in terms of the symmetries and the number of linear regions that they have. Deep networks are able to sequentially map portions of each layer's input-space to the same output. In this way, deep models compute functions that react equally to complicated patterns of different inputs. The compositional structure of these functions enables them to re-use pieces of computation exponentially often in terms of the network's depth. This paper investigates the complexity of such compositional maps and contributes new theoretical results regarding the advantage of depth for neural networks with piecewise linear activation functions. In particular, our analysis is not specific to a single family of models, and as an example, we employ it for rectifier and maxout networks. We improve complexity bounds from pre-existing work and investigate the behavior of units in higher layers.
References in corpus (2)
Cited by in corpus (258)
- Deep Residual Learning for Image Recognition
- Methods for Interpreting and Understanding Deep Neural Networks
- MobileNetV2: Inverted Residuals and Linear Bottlenecks
- Wide Residual Networks
- Explaining Deep Neural Networks and Beyond: A Review of Methods and Applications
- Training Very Deep Networks
- Deep Learning for Single Image Super-Resolution: A Brief Review
- Impact of Fully Connected Layers on Performance of Convolutional Neural Networks for Image Classification
- Optimization Problems for Machine Learning: A Survey
- Efficient representation and approximation of model predictive control laws via deep learning
- Breaking the Curse of Dimensionality with Convex Neural Networks
- Data Driven Governing Equations Approximation Using Deep Neural Networks
- Nonparametric regression using deep neural networks with ReLU activation function
- Exponential expressivity in deep neural networks through transient chaos
- Sensitivity and Generalization in Neural Networks: an Empirical Study
- Discovering Symbolic Models from Deep Learning with Inductive Biases
- Robust Large Margin Deep Neural Networks
- Deep Residual Learning in Spiking Neural Networks
- The Power of Depth for Feedforward Neural Networks
- Provable approximation properties for deep neural networks
- Generalization in Deep Learning
- Deep Network Approximation for Smooth Functions
- Deep Neural Network Compression for Aircraft Collision Avoidance Systems
- Bridging the Gaps Between Residual Learning, Recurrent Neural Networks and Visual Cortex
- Essentially No Barriers in Neural Network Energy Landscape
- Deep Neural Networks with Random Gaussian Weights: A Universal Classification Strategy?
- A Theory of Generative ConvNet
- ReZero is All You Need: Fast Convergence at Large Depth
- Sharp Minima Can Generalize For Deep Nets
- The Modern Mathematics of Deep Learning
- The jamming transition as a paradigm to understand the loss landscape of deep neural networks
- The Convergence Rate of Neural Networks for Learned Functions of Different Frequencies
- Deep learning generalizes because the parameter-function map is biased towards simple functions
- Are All Layers Created Equal?
- Learning Functions: When Is Deep Better Than Shallow
- The Shattered Gradients Problem: If resnets are the answer, then what is the question?
- Fermionic Wave Functions from Neural-Network Constrained Hidden States
- Autonomous Deep Learning: Continual Learning Approach for Dynamic Environments
- Optimal Approximation Rate of ReLU Networks in terms of Width and Depth
- Batch-normalized Maxout Network in Network
- Neural Models for Information Retrieval
- Nonlinear Approximation via Compositions
- Multifidelity Modeling for Physics-Informed Neural Networks (PINNs)
- Sorting out Lipschitz function approximation
- Depth with Nonlinearity Creates No Bad Local Minima in ResNets
- Knowledge Transfer via Distillation of Activation Boundaries Formed by Hidden Neurons
- Soft-Demapping for Short Reach Optical Communication: A Comparison of Deep Neural Networks and Volterra Series
- Stiffness: A New Perspective on Generalization in Neural Networks
- Normalization Techniques in Training DNNs: Methodology, Analysis and Application
- Deep Learning Based Packet Detection and Carrier Frequency Offset Estimation in IEEE 802.11ah
- Advances in Quantum Deep Learning: An Overview
- Weisfeiler and Lehman Go Topological: Message Passing Simplicial Networks
- Universal Statistics of Fisher Information in Deep Neural Networks: Mean Field Approach
- K for the Price of 1: Parameter-efficient Multi-task and Transfer Learning
- Bounding and Counting Linear Regions of Deep Neural Networks
- Effect of Depth and Width on Local Minima in Deep Learning
- On the Selection of Initialization and Activation Function for Deep Neural Networks
- Towards Understanding Sparse Filtering: A Theoretical Perspective
- Reliably-stabilizing piecewise-affine neural network controllers
- What Kinds of Functions do Deep Neural Networks Learn? Insights from Variational Spline Theory
- Competitive Multi-scale Convolution
- More is Less: A More Complicated Network with Less Inference Complexity
- Robust Counterfactual Explanations on Graph Neural Networks
- Expressiveness of Rectifier Networks
- ReLU Deep Neural Networks from the Hierarchical Basis Perspective
- DecomposeMe: Simplifying ConvNets for End-to-End Learning
- Deep Neural Networks with Trainable Activations and Controlled Lipschitz Constant
- DeepOPF: A Deep Neural Network Approach for Security-Constrained DC Optimal Power Flow
- Tropical Geometry of Deep Neural Networks
- A Deep Learning Approach to Unsupervised Ensemble Learning
- Physical deep learning based on optimal control of dynamical systems
- Vulnerabilities of Connectionist AI Applications: Evaluation and Defence
- Empirical Studies on the Properties of Linear Regions in Deep Neural Networks
- On the Number of Linear Regions of Convolutional Neural Networks
- Liberty or Depth: Deep Bayesian Neural Nets Do Not Need Complex Weight Posterior Approximations
- CTU Depth Decision Algorithms for HEVC: A Survey
- A convolutional neural network approach for reconstructing polarization information of photoelectric X-ray polarimeters
- Deep learning is adaptive to intrinsic dimensionality of model smoothness in anisotropic Besov space
- Fast Neural Network Verification via Shadow Prices
- Deep-RBF Networks Revisited: Robust Classification with Rejection
- Approximation Properties of Deep ReLU CNNs
- MAMNet: Multi-path Adaptive Modulation Network for Image Super-Resolution
- A Neural Scaling Law from the Dimension of the Data Manifold
- Unwrapping The Black Box of Deep ReLU Networks: Interpretability, Diagnostics, and Simplification
- Theoretical Issues in Deep Networks: Approximation, Optimization and Generalization
- Learning Deep Analysis Dictionaries for Image Super-Resolution
- A RAD approach to deep mixture models
- Efficient Representation of Low-Dimensional Manifolds using Deep Networks
- Is Deeper Better only when Shallow is Good?
- Neural Networks Should Be Wide Enough to Learn Disconnected Decision Regions
- Some Theoretical Insights into Wasserstein GANs
- Visualizing the PHATE of Neural Networks
- FFT Convolutions are Faster than Winograd on Modern CPUs, Here is Why
- Butterfly-Net: Optimal Function Representation Based on Convolutional Neural Networks
- Discovering and Explaining the Representation Bottleneck of DNNs
- Towards Robust, Locally Linear Deep Networks
- Do Neural Nets Learn Statistical Laws behind Natural Language?
- A Probabilistic Framework for Deep Learning
- Exact and Consistent Interpretation for Piecewise Linear Neural Networks: A Closed Form Solution
- Uncertainty Quantification for Sparse Deep Learning
- How Graph Neural Network Interatomic Potentials Extrapolate: Role of the Message-Passing Algorithm
- Deep Active Learning for Text Classification with Diverse Interpretations
- Three-dimensional convolutional neural networks for neutrinoless double-beta decay signal/background discrimination in high-pressure gaseous Time Projection Chamber
- Approximation in shift-invariant spaces with deep ReLU neural networks
- CKConv: Continuous Kernel Convolution For Sequential Data
- Provably Good Solutions to the Knapsack Problem via Neural Networks of Bounded Size
- When are Deep Networks really better than Decision Forests at small sample sizes, and how?
- Artificial Intelligence Assisted Inversion (AIAI) of Synthetic Type Ia Supernova Spectra
- Zen-NAS: A Zero-Shot NAS for High-Performance Deep Image Recognition
- Variational Laplace Autoencoders
- Development of Use-specific High Performance Cyber-Nanomaterial Optical Detectors by Effective Choice of Machine Learning Algorithms
- Lightweight and Efficient Image Super-Resolution with Block State-based Recursive Network
- Reachability Analysis for Feed-Forward Neural Networks using Face Lattices
- Equivalent and Approximate Transformations of Deep Neural Networks
- On Arrhythmia Detection by Deep Learning and Multidimensional Representation
- Compression based bound for non-compressed network: unified generalization error analysis of large compressible deep neural network
- Dissecting Deep Neural Networks
- A closer look at the approximation capabilities of neural networks
- Synthesizing Dynamic Patterns by Spatial-Temporal Generative ConvNet
- The Local Elasticity of Neural Networks
- Deep Learning for MIMO Channel Estimation: Interpretation, Performance, and Comparison
- Smooth activations and reproducibility in deep networks
- Training Neural Networks by Using Power Linear Units (PoLUs)
- Subspace Preserving Quantum Convolutional Neural Network Architectures
- Orthogonal Deep Neural Networks
- Analysis on the Nonlinear Dynamics of Deep Neural Networks: Topological Entropy and Chaos
- The Nonlinearity Coefficient - Predicting Generalization in Deep Neural Networks
- Optimal Nonparametric Inference via Deep Neural Network
- Regularizing Towards Permutation Invariance in Recurrent Models
- The Shallow End: Empowering Shallower Deep-Convolutional Networks through Auxiliary Outputs
- Tropical Polynomial Division and Neural Networks
- Data Augmentation for Bayesian Deep Learning
- Analysis of Invariance and Robustness via Invertibility of ReLU-Networks
- Deep ReLU Networks Preserve Expected Length
- On the Depth of Deep Neural Networks: A Theoretical View
- Data-Driven Robust Optimization using Unsupervised Deep Learning
- Depth-Width Trade-offs for Neural Networks via Topological Entropy
- Model Complexity of Deep Learning: A Survey
- Conditional Computation for Continual Learning
- Understanding Deep MIMO Detection
- Learning Strict Identity Mappings in Deep Residual Networks
- Frequency-compensated PINNs for Fluid-dynamic Design Problems
- Topology of deep neural networks
- Hierarchical Decomposition of Nonlinear Dynamics and Control for System Identification and Policy Distillation
- An Analysis of the Expressiveness of Deep Neural Network Architectures Based on Their Lipschitz Constants
- Representational Capacity of Deep Neural Networks -- A Computing Study
- Adaptive Reconstruction Network for Weakly Supervised Referring Expression Grounding
- Enhancing Adversarial Defense by k-Winners-Take-All
- Numerically Solving Parametric Families of High-Dimensional Kolmogorov Partial Differential Equations via Deep Learning
- On the Compressive Power of Boolean Threshold Autoencoders
- The Limitations of Large Width in Neural Networks: A Deep Gaussian Process Perspective
- Optimal Function Approximation with Relu Neural Networks
- Self-Supervised Deep Visual Odometry with Online Adaptation
- Optimising for Interpretability: Convolutional Dynamic Alignment Networks
- Approximating Lipschitz continuous functions with GroupSort neural networks
- Learning from Learning Machines: Optimisation, Rules, and Social Norms
- Automatic Construction of Multi-layer Perceptron Network from Streaming Examples
- Measuring Model Complexity of Neural Networks with Curve Activation Functions
- An Embedding of ReLU Networks and an Analysis of their Identifiability
- Towards Lower Bounds on the Depth of ReLU Neural Networks
- Sharp bounds for the number of regions of maxout networks and vertices of Minkowski sums
- Scaling Up Exact Neural Network Compression by ReLU Stability
- DNN2LR: Interpretation-inspired Feature Crossing for Real-world Tabular Data
- Hierarchically Compositional Tasks and Deep Convolutional Networks
- Function approximation by deep networks
- Deep Fundamental Factor Models
- Nearly-tight bounds on linear regions of piecewise linear neural networks
- A representer theorem for deep neural networks
- On Deep Representation Learning from Noisy Web Images
- Compressive MRI quantification using convex spatiotemporal priors and deep auto-encoders
- Geometry of Deep Convolutional Networks
- Traversing the Local Polytopes of ReLU Neural Networks: A Unified Approach for Network Verification
- A Review of Formal Methods applied to Machine Learning
- An Efficient Explorative Sampling Considering the Generative Boundaries of Deep Generative Neural Networks
- How the brain might work: statistics flowing in redundant population codes
- Tailored neural networks for learning optimal value functions in MPC
- On the Size and Width of the Decoder of a Boolean Threshold Autoencoder
- Debona: Decoupled Boundary Network Analysis for Tighter Bounds and Faster Adversarial Robustness Proofs
- Better Depth-Width Trade-offs for Neural Networks through the lens of Dynamical Systems
- What training reveals about neural network complexity
- Reachability Analysis of Convolutional Neural Networks
- Knowledge-guided Pairwise Reconstruction Network for Weakly Supervised Referring Expression Grounding
- Neuron ranking -- an informed way to condense convolutional neural networks architecture
- Truncating Wide Networks using Binary Tree Architectures
- Training Feedforward Neural Networks with Standard Logistic Activations is Feasible
- Theoretical Analysis of the Advantage of Deepening Neural Networks
- Depth separation beyond radial functions
- The Geometry of Deep Networks: Power Diagram Subdivision
- Analyzing Finite Neural Networks: Can We Trust Neural Tangent Kernel Theory?
- When Hardness of Approximation Meets Hardness of Learning
- Feature selection of neural networks is skewed towards the less abstract cue
- MaGNET: Uniform Sampling from Deep Generative Network Manifolds Without Retraining
- On the Number of Linear Functions Composing Deep Neural Network: Towards a Refined Definition of Neural Networks Complexity
- Investigating the Compositional Structure Of Deep Neural Networks
- Fast gradient-free activation maximization for neurons in spiking neural networks
- Can Deep Learning Predict Risky Retail Investors? A Case Study in Financial Risk Behavior Forecasting
- A partition-based similarity for classification distributions
- Neural networks behave as hash encoders: An empirical study
- Learning Energy-Based Models as Generative ConvNets via Multi-grid Modeling and Sampling
- Min-Max-Plus Neural Networks
- Tight Bounds on the Smallest Eigenvalue of the Neural Tangent Kernel for Deep ReLU Networks
- Using activation histograms to bound the number of affine regions in ReLU feed-forward neural networks
- Deep Unitary Convolutional Neural Networks
- Universal Approximation of Residual Flows in Maximum Mean Discrepancy
- DWDN: Deep Wiener Deconvolution Network for Non-Blind Image Deblurring
- Bounding The Number of Linear Regions in Local Area for Neural Networks with ReLU Activations
- A Multivariate Density Forecast Approach for Online Power System Security Assessment
- Learning Energy-based Spatial-Temporal Generative ConvNets for Dynamic Patterns
- Deep Stacked Stochastic Configuration Networks for Lifelong Learning of Non-Stationary Data Streams
- Deep Reinforcement Learning with Linear Quadratic Regulator Regions
- Sampling Using Neural Networks for colorizing the grayscale images
- Test Sample Accuracy Scales with Training Sample Density in Neural Networks
- Shallow Neural Network can Perfectly Classify an Object following Separable Probability Distribution
- In Proximity of ReLU DNN, PWA Function, and Explicit MPC
- Fast generalization error bound of deep learning without scale invariance of activation functions
- A Tour of Convolutional Networks Guided by Linear Interpreters
- Limiting Network Size within Finite Bounds for Optimization
- Collocation approximation by deep neural ReLU networks for parametric elliptic PDEs with lognormal inputs
- Complexity for deep neural networks and other characteristics of deep feature representations
- Learning Boolean functions with concentrated spectra
- Knots in random neural networks
- Deep ReLU Programming
- Exact and Consistent Interpretation of Piecewise Linear Models Hidden behind APIs: A Closed Form Solution
- Interpretations of Deep Learning by Forests and Haar Wavelets
- Random Bias Initialization Improves Quantized Training
- Robust Deep Neural Networks Inspired by Fuzzy Logic
- Decomposing tropical rational functions
- Automatic Discoveries of Physical and Semantic Concepts via Association Priors of Neuron Groups
- Probabilistic bounds on neuron death in deep rectifier networks
- Group Whitening: Balancing Learning Efficiency and Representational Capacity
- Deep Autoencoders: From Understanding to Generalization Guarantees
- Understanding Interpretability by generalized distillation in Supervised Classification
- Network Implosion: Effective Model Compression for ResNets via Static Layer Pruning and Retraining
- On the Expected Complexity of Maxout Networks
- How Could Polyhedral Theory Harness Deep Learning?
- Expressive Power and Loss Surfaces of Deep Learning Models
- Reducing Parameter Space for Neural Network Training
- Training BatchNorm Only in Neural Architecture Search and Beyond
- How Analysis Can Teach Us the Optimal Way to Design Neural Operators
- Unsupervised Representation Learning via Neural Activation Coding
- Expressivity of Neural Networks via Chaotic Itineraries beyond Sharkovsky's Theorem
- Efficient and Robust Mixed-Integer Optimization Methods for Training Binarized Deep Neural Networks
- The Self-Simplifying Machine: Exploiting the Structure of Piecewise Linear Neural Networks to Create Interpretable Models
- Trainable back-propagated functional transfer matrices
- DNN2LR: Automatic Feature Crossing for Credit Scoring
- Spurious Local Minima Are Common for Deep Neural Networks with Piecewise Linear Activations
- Fast Jacobian-Vector Product for Deep Networks
- Bias-corrected estimator for intrinsic dimension and differential entropy--a visual multiscale approach
- Goal Recognition over Imperfect Domain Models
- ReLU Code Space: A Basis for Rating Network Quality Besides Accuracy
- The Restricted Isometry of ReLU Networks: Generalization through Norm Concentration
- Deep Active Learning by Model Interpretability
- A Partial Regularization Method for Network Compression
- Towards Quantifying Intrinsic Generalization of Deep ReLU Networks
- Deep Representation with ReLU Neural Networks
- Tucker Decomposition Network: Expressive Power and Comparison
- Light Multi-segment Activation for Model Compression
- The many faces of deep learning