Opening the Black Box of Deep Neural Networks via Information
arXiv:1703.00810
Abstract
Despite their great success, there is still no comprehensive theoretical understanding of learning with Deep Neural Networks (DNNs) or their inner organization. Previous work proposed to analyze DNNs in the \textit{Information Plane}; i.e., the plane of the Mutual Information values that each layer preserves on the input and output variables. They suggested that the goal of the network is to optimize the Information Bottleneck (IB) tradeoff between compression and prediction, successively, for each layer. In this work we follow up on this idea and demonstrate the effectiveness of the Information-Plane visualization of DNNs. Our main results are: (i) most of the training epochs in standard DL are spent on {\emph compression} of the input to efficient representation and not on fitting the training labels. (ii) The representation compression phase begins when the training errors becomes small and the Stochastic Gradient Decent (SGD) epochs change from a fast drift to smaller training error into a stochastic relaxation, or random diffusion, constrained by the training error value. (iii) The converged layers lie on or very close to the Information Bottleneck (IB) theoretical bound, and the maps from the input to any hidden layer and from this hidden layer to the output satisfy the IB self-consistent equations. This generalization through noise mechanism is unique to Deep Neural Networks and absent in one layer networks. (iv) The training time is dramatically reduced when adding more hidden layers. Thus the main advantage of the hidden layers is computational. This can be explained by the reduced relaxation time, as this it scales super-linearly (exponentially for simple diffusion) with the information compression from the previous layer.
19 pages, 8 figures
References in corpus (2)
Cited by in corpus (325)
- Machine learning and the physical sciences
- A Survey on Explainable Artificial Intelligence (XAI): Towards Medical XAI
- Physics Informed Deep Learning (Part I): Data-driven Solutions of Nonlinear Partial Differential Equations
- When Does Label Smoothing Help?
- Discovering physical concepts with neural networks
- Generalisation in humans and deep neural networks
- A Survey on Data Augmentation for Text Classification
- Artificial neural networks for neuroscientists: A primer
- NVAE: A Deep Hierarchical Variational Autoencoder
- Dimensionality-Driven Learning with Noisy Labels
- On the importance of single directions for generalization
- Three Factors Influencing Minima in SGD
- Explanation in Human-AI Systems: A Literature Meta-Review, Synopsis of Key Ideas and Publications, and Bibliography for Explainable AI
- Adversarial Examples: Attacks and Defenses for Deep Learning
- Explainable AI: Beware of Inmates Running the Asylum Or: How I Learnt to Stop Worrying and Love the Social and Behavioural Sciences
- Optimal Errors and Phase Transitions in High-Dimensional Generalized Linear Models
- Learning Representations for Neural Network-Based Classification Using the Information Bottleneck Principle
- Nonlinear Information Bottleneck
- Intrinsic dimension of data representations in deep neural networks
- Beyond Finite Layer Neural Networks: Bridging Deep Architectures and Numerical Differential Equations
- Machine Learning of Explicit Order Parameters: From the Ising Model to SU(2) Lattice Gauge Theory
- Empirical Analysis of the Hessian of Over-Parametrized Neural Networks
- Entropy and mutual information in models of deep neural networks
- An Empirical Model of Large-Batch Training
- Approximating Continuous Functions by ReLU Nets of Minimal Width
- Symmetric Cross Entropy for Robust Learning with Noisy Labels
- The jamming transition as a paradigm to understand the loss landscape of deep neural networks
- A Deep Information Sharing Network for Multi-contrast Compressed Sensing MRI Reconstruction
- Estimation of energy consumption of electric vehicles using Deep Convolutional Neural Network to reduce driver's range anxiety
- Adapting Auxiliary Losses Using Gradient Similarity
- How to Certify Machine Learning Based Safety-critical Systems? A Systematic Literature Review
- Mathematics of Deep Learning
- Revisiting Training Strategies and Generalization Performance in Deep Metric Learning
- On Identifiability in Transformers
- Polynomial Regression As an Alternative to Neural Nets
- Predicting trends in the quality of state-of-the-art neural networks without access to training or testing data
- Deep learning as a tool for neural data analysis: speech classification and cross-frequency coupling in human sensorimotor cortex
- Information Scrambling in Quantum Neural Networks
- Critical Learning Periods in Deep Neural Networks
- Neural Persistence: A Complexity Measure for Deep Neural Networks Using Algebraic Topology
- Credit spread approximation and improvement using random forest regression
- Excessive Invariance Causes Adversarial Vulnerability
- i-RevNet: Deep Invertible Networks
- A Survey of Safety and Trustworthiness of Deep Neural Networks: Verification, Testing, Adversarial Attack and Defence, and Interpretability
- Batch Normalization is a Cause of Adversarial Vulnerability
- Automated Pruning for Deep Neural Network Compression
- On Information Plane Analyses of Neural Network Classifiers -- A Review
- What shapes feature representations? Exploring datasets, architectures, and training
- Statistical Criticality arises in Most Informative Representations
- Inspect, Understand, Overcome: A Survey of Practical Methods for AI Safety
- SGD on Neural Networks Learns Functions of Increasing Complexity
- Training behavior of deep neural network in frequency domain
- Deep Learning Theory Review: An Optimal Control and Dynamical Systems Perspective
- 2PFPCE: Two-Phase Filter Pruning Based on Conditional Entropy
- Where is the Information in a Deep Neural Network?
- On Interpretability of Artificial Neural Networks: A Survey
- Are Sixteen Heads Really Better than One?
- Towards a mathematical framework to inform Neural Network modelling via Polynomial Regression
- Fixing a Broken ELBO
- Mean Field Limit of the Learning Dynamics of Multilayer Neural Networks
- A Comprehensive Survey of Machine Learning Applied to Radar Signal Processing
- Computing Quantum Channel Capacities
- Honest-but-Curious Nets: Sensitive Attributes of Private Inputs Can Be Secretly Coded into the Classifiers' Outputs
- InfoGCL: Information-Aware Graph Contrastive Learning
- On Universal Features for High-Dimensional Learning and Inference
- How higher goals are constructed and collapse under stress: a hierarchical Bayesian control systems perspective
- Multi-modal Deep Guided Filtering for Comprehensible Medical Image Processing
- Image classification using quantum inference on the D-Wave 2X
- Auto-Encoding Total Correlation Explanation
- Stochastic Mirror Descent on Overparameterized Nonlinear Models: Convergence, Implicit Regularization, and Generalization
- Dual Semantic Fusion Network for Video Object Detection
- Parameterized Reinforcement Learning for Optical System Optimization
- Deep Semi-Supervised Anomaly Detection
- Information Theoretic Counterfactual Learning from Missing-Not-At-Random Feedback
- Understanding Neural Networks and Individual Neuron Importance via Information-Ordered Cumulative Ablation
- Convexity and Operational Interpretation of the Quantum Information Bottleneck Function
- Spatiotemporal Filtering for Event-Based Action Recognition
- Deep Learning on Image Denoising: An overview
- Neural network representation of electronic structure from molecular dynamics
- Interpolated Adversarial Training: Achieving Robust Neural Networks without Sacrificing Too Much Accuracy
- Rethinking generalization requires revisiting old ideas: statistical mechanics approaches and complex learning behavior
- The RFML Ecosystem: A Look at the Unique Challenges of Applying Deep Learning to Radio Frequency Applications
- Revisiting Locally Supervised Learning: an Alternative to End-to-end Training
- Measuring Compositionality in Representation Learning
- Beyond accuracy: quantifying trial-by-trial behaviour of CNNs and humans by measuring error consistency
- Gradient Descent Quantizes ReLU Network Features
- Information Plane Analysis of Deep Neural Networks via Matrix-Based Renyi's Entropy and Tensor Kernels
- Physics Informed Deep Learning for Transport in Porous Media. Buckley Leverett Problem
- Reducing Information Bottleneck for Weakly Supervised Semantic Segmentation
- The Deep Kernelized Autoencoder
- Removing Bias in Multi-modal Classifiers: Regularization by Maximizing Functional Entropies
- Knowledge Consistency between Neural Networks and Beyond
- Estimating the Mutual Information between two Discrete, Asymmetric Variables with Limited Samples
- Learning Optimal Representations with the Decodable Information Bottleneck
- Caveats for information bottleneck in deterministic scenarios
- High-Fidelity GAN Inversion for Image Attribute Editing
- Unsupervised Learning of Neural Networks to Explain Neural Networks
- Kernelized information bottleneck leads to biologically plausible 3-factor Hebbian learning in deep networks
- Generating Weighted MAX-2-SAT Instances of Tunable Difficulty with Frustrated Loops
- Technical Considerations for Semantic Segmentation in MRI using Convolutional Neural Networks
- The Bottom-up Evolution of Representations in the Transformer: A Study with Machine Translation and Language Modeling Objectives
- LCA: Loss Change Allocation for Neural Network Training
- Explaining Knowledge Distillation by Quantifying the Knowledge
- On Data-Augmentation and Consistency-Based Semi-Supervised Learning
- Discovering and Explaining the Representation Bottleneck of DNNs
- Confidential Inference via Ternary Model Partitioning
- Quantifying the generalization error in deep learning in terms of data distribution and neural network smoothness
- Dynamic Bottleneck for Robust Self-Supervised Exploration
- Explaining a black-box using Deep Variational Information Bottleneck Approach
- Training Invertible Neural Networks as Autoencoders
- Information flows of diverse autoencoders
- Layer-wise Characterization of Latent Information Leakage in Federated Learning
- "Dependency Bottleneck" in Auto-encoding Architectures: an Empirical Study
- Understanding Convolutional Neural Networks with Information Theory: An Initial Exploration
- The Dual Information Bottleneck
- On the Dimensionality of Embeddings for Sparse Features and Data
- On Network Science and Mutual Information for Explaining Deep Neural Networks
- Information Bottleneck and its Applications in Deep Learning
- Farewell to Mutual Information: Variational Distillation for Cross-Modal Person Re-Identification
- Learning to Learn with Variational Information Bottleneck for Domain Generalization
- An Information Theory-inspired Strategy for Automatic Network Pruning
- Ablation of a Robot's Brain: Neural Networks Under a Knife
- Inflation as an Information Bottleneck - A strategy for identifying universality classes and making robust predictions
- PAC-Bayes Information Bottleneck
- Bounded Information Rate Variational Autoencoders
- Tight Mutual Information Estimation With Contrastive Fenchel-Legendre Optimization
- Categorical Perception: A Groundwork for Deep Learning
- FRODO: Free rejection of out-of-distribution samples: application to chest x-ray analysis
- Fundamental Limits and Tradeoffs in Invariant Representation Learning
- Efficient human-like semantic representations via the Information Bottleneck principle
- Compression-Based Regularization with an Application to Multi-Task Learning
- On the Information Plane of Autoencoders
- Quasi-potential as an implicit regularizer for the loss function in the stochastic gradient descent
- Expressive power of recurrent neural networks
- Invariance of Weight Distributions in Rectified MLPs
- Unpacking Information Bottlenecks: Unifying Information-Theoretic Objectives in Deep Learning
- Mutual Information Scaling and Expressive Power of Sequence Models
- Layer-wise Learning of Stochastic Neural Networks with Information Bottleneck
- Bioinformatics and Medicine in the Era of Deep Learning
- Filter Grafting for Deep Neural Networks
- Sliced Mutual Information: A Scalable Measure of Statistical Dependence
- Information Losses in Neural Classifiers from Sampling
- Information in Infinite Ensembles of Infinitely-Wide Neural Networks
- Usable Information and Evolution of Optimal Representations During Training
- On the Maximum Mutual Information Capacity of Neural Architectures
- Is SGD a Bayesian sampler? Well, almost
- Revisiting Hilbert-Schmidt Information Bottleneck for Adversarial Robustness
- DisCo: Remedy Self-supervised Learning on Lightweight Models with Distilled Contrastive Learning
- Estimating Differential Entropy under Gaussian Convolutions
- Renormalized Mutual Information for Artificial Scientific Discovery
- ToyArchitecture: Unsupervised Learning of Interpretable Models of the World
- Neuron Campaign for Initialization Guided by Information Bottleneck Theory
- Mutual Information Gradient Estimation for Representation Learning
- Information Bottleneck Theory on Convolutional Neural Networks
- Measuring Dependence with Matrix-based Entropy Functional
- Improving Robustness to Model Inversion Attacks via Mutual Information Regularization
- Gaussian Mixture Models for Blended Photometric Redshifts
- Differentiable programming and its applications to dynamical systems
- Relative stability toward diffeomorphisms indicates performance in deep nets
- Dynamic learning rate using Mutual Information
- Training Normalizing Flows with the Information Bottleneck for Competitive Generative Classification
- Reconstruction Bottlenecks in Object-Centric Generative Models
- Deep Deterministic Information Bottleneck with Matrix-based Entropy Functional
- On the Effect of Low-Rank Weights on Adversarial Robustness of Neural Networks
- The Role of Information Complexity and Randomization in Representation Learning
- Partially Observable Szilard Engines
- Designing Artificial Cognitive Architectures: Brain Inspired or Biologically Inspired?
- Server, server in the cloud. Who is the fairest in the crowd?
- Understanding Deep Learning Generalization by Maximum Entropy
- Gaussian Lower Bound for the Information Bottleneck Limit
- Why ResNet Works? Residuals Generalize
- Improving Robustness and Generality of NLP Models Using Disentangled Representations
- Partial local entropy and anisotropy in deep weight spaces
- Multiscale Principle of Relevant Information for Hyperspectral Image Classification
- Shredder: Learning Noise Distributions to Protect Inference Privacy
- Fast Convergence for Langevin Diffusion with Manifold Structure
- Informative Neural Ensemble Kalman Learning
- The Variational Deficiency Bottleneck
- Focus of Attention Improves Information Transfer in Visual Features
- Spontaneous Symmetry Breaking in Neural Networks
- Translating Diffusion, Wavelets, and Regularisation into Residual Networks
- Deep Dimension Reduction for Supervised Representation Learning
- Perspective: A Phase Diagram for Deep Learning unifying Jamming, Feature Learning and Lazy Training
- Trust but Verify: An Information-Theoretic Explanation for the Adversarial Fragility of Machine Learning Systems, and a General Defense against Adversarial Attacks
- The Dimpled Manifold Model of Adversarial Examples in Machine Learning
- General Information Bottleneck Objectives and their Applications to Machine Learning
- Critical Slowing Down Near Topological Transitions in Rate-Distortion Problems
- Towards a theory of machine learning
- On the Estimation of Information Measures of Continuous Distributions
- Edge of chaos as a guiding principle for modern neural network training
- Contrasting information theoretic decompositions of modulatory and arithmetic interactions in neural information processing systems
- What Information Does a ResNet Compress?
- Mixup Regularization for Region Proposal based Object Detectors
- Abstraction Mechanisms Predict Generalization in Deep Neural Networks
- Power System Event Identification based on Deep Neural Network with Information Loading
- Explaining Natural Language Processing Classifiers with Occlusion and Language Modeling
- Stochastic gradient descent with random learning rate
- Modeling Information Flow Through Deep Neural Networks
- PRI-VAE: Principle-of-Relevant-Information Variational Autoencoders
- UFANS: U-shaped Fully-Parallel Acoustic Neural Structure For Statistical Parametric Speech Synthesis With 20X Faster
- Collective evolution of weights in wide neural networks
- Bregman Divergence Bounds and Universality Properties of the Logarithmic Loss
- A synthetic dataset for deep learning
- Restricted Boltzmann Machine Flows and The Critical Temperature of Ising models
- Explaining AlphaGo: Interpreting Contextual Effects in Neural Networks
- Scaling Object Detection by Transferring Classification Weights
- Spectral Roll-off Points Variations: Exploring Useful Information in Feature Maps by Its Variations
- Learning Representations in Reinforcement Learning:An Information Bottleneck Approach
- Multi-Lead ECG Classification via an Information-Based Attention Convolutional Neural Network
- A Simple yet Effective Baseline for Robust Deep Learning with Noisy Labels
- Wider Networks Learn Better Features
- Data augmentation and image understanding
- Capacity-Approaching Autoencoders for Communications
- Drill the Cork of Information Bottleneck by Inputting the Most Important Data
- A Free-Energy Principle for Representation Learning
- A Probabilistic Representation of DNNs: Bridging Mutual Information and Generalization
- Do Compressed Representations Generalize Better?
- Interpreting and Disentangling Feature Components of Various Complexity from DNNs
- Learning to Find Correlated Features by Maximizing Information Flow in Convolutional Neural Networks
- Hierarchical nucleation in deep neural networks
- To Beta or Not To Beta: Information Bottleneck for DigitaL Image Forensics
- Feature selection of neural networks is skewed towards the less abstract cue
- Understanding Learning Dynamics for Neural Machine Translation
- Representation Edit Distance as a Measure of Novelty
- A Differential Game Theoretic Neural Optimizer for Training Residual Networks
- Information Theoretic Interpretation of Deep learning
- The Helmholtz Method: Using Perceptual Compression to Reduce Machine Learning Complexity
- Deep Epitome for Unravelling Generalized Hamming Network: A Fuzzy Logic Interpretation of Deep Learning
- Gradient Normalization & Depth Based Decay For Deep Learning
- Information Bottleneck: Exact Analysis of (Quantized) Neural Networks
- A Probabilistic Representation of Deep Learning for Improving The Information Theoretic Interpretability
- Evaluation of Dataflow through layers of Deep Neural Networks in Classification and Regression Problems
- Learning View-Disentangled Human Pose Representation by Contrastive Cross-View Mutual Information Maximization
- Trap of Feature Diversity in the Learning of MLPs
- Information Plane Analysis Visualization in Deep Learning via Transfer Entropy
- Understanding Learning Dynamics of Binary Neural Networks via Information Bottleneck
- Estimating informativeness of samples with Smooth Unique Information
- Understanding the Effects of Pre-Training for Object Detectors via Eigenspectrum
- The role of a layer in deep neural networks: a Gaussian Process perspective
- Why Do Better Loss Functions Lead to Less Transferable Features?
- Simplifying the explanation of deep neural networks with sufficient and necessary feature-sets: case of text classification
- An Information-theoretic Visual Analysis Framework for Convolutional Neural Networks
- Information Bottleneck for an Oblivious Relay with Channel State Information: the Vector Case
- Properties of the After Kernel
- InfoNEAT: Information Theory-based NeuroEvolution of Augmenting Topologies for Side-channel Analysis
- Understanding Neural Networks with Logarithm Determinant Entropy Estimator
- AKE-GNN: Effective Graph Learning with Adaptive Knowledge Exchange
- A Probabilistic Representation of Deep Learning
- Inference-InfoGAN: Inference Independence via Embedding Orthogonal Basis Expansion
- ClusterNet: Detecting Small Objects in Large Scenes by Exploiting Spatio-Temporal Information
- Generalized Constraints as A New Mathematical Problem in Artificial Intelligence: A Review and Perspective
- Separation of time scales and direct computation of weights in deep neural networks
- Phases of learning dynamics in artificial neural networks: with or without mislabeled data
- Towards Safety Verification of Direct Perception Neural Networks
- Blocked and Hierarchical Disentangled Representation From Information Theory Perspective
- Analyzing Data Selection Techniques with Tools from the Theory of Information Losses
- The distance between the weights of the neural network is meaningful
- Generalisation in fully-connected neural networks for time series forecasting
- A Framework for Behavioral Biometric Authentication using Deep Metric Learning on Mobile Devices
- Measure, Manifold, Learning, and Optimization: A Theory Of Neural Networks
- Bayesian Convolutional Neural Networks for Compressed Sensing Restoration
- Ultra-light deep MIR by trimming lottery tickets
- Bayesian deep learning for mapping via auxiliary information: a new era for geostatistics?
- An Extension of Fano's Inequality for Characterizing Model Susceptibility to Membership Inference Attacks
- Faster Convergence & Generalization in DNNs
- Mean Field Theory of Activation Functions in Deep Neural Networks
- Cortex Neural Network: learning with Neural Network groups
- Information-Theoretic Perspective of Federated Learning
- Dropping Networks for Transfer Learning
- Privacy for Rescue: A New Testimony Why Privacy is Vulnerable In Deep Models
- Margin Maximization as Lossless Maximal Compression
- Explainable Deep RDFS Reasoner
- Generalization of an Upper Bound on the Number of Nodes Needed to Achieve Linear Separability
- ReNN: Rule-embedded Neural Networks
- An Information-Theoretic Explanation for the Adversarial Fragility of AI Classifiers
- Information Theoretic Lower Bounds on Negative Log Likelihood
- Resolution and Relevance Trade-offs in Deep Learning
- Scientists in silico?
- Seeing Convolution Through the Eyes of Finite Transformation Semigroup Theory: An Abstract Algebraic Interpretation of Convolutional Neural Networks
- Identification of state functions by physically-guided neural networks with physically-meaningful internal layers
- Towards Further Understanding of Sparse Filtering via Information Bottleneck
- Examining the causal structures of deep neural networks using information theory
- Efficient decorrelation of features using Gramian in Reinforcement Learning
- Understanding the Behaviour of the Empirical Cross-Entropy Beyond the Training Distribution
- A Practical & Unified Notation for Information-Theoretic Quantities in ML
- A Generalization Theory based on Independent and Task-Identically Distributed Assumption
- Sparsity Emerges Naturally in Neural Language Models
- Improve variational autoEncoder with auxiliary softmax multiclassifier
- Sparsity-Probe: Analysis tool for Deep Learning Models
- Partial Or Complete, That's The Question
- Topologically Densified Distributions
- DNNs as Layers of Cooperating Classifiers
- What are Neural Networks made of?
- From abstract items to latent spaces to observed data and back: Compositional Variational Auto-Encoder
- Function space analysis of deep learning representation layers
- Detecting Learning vs Memorization in Deep Neural Networks using Shared Structure Validation Sets
- Is the Meta-Learning Idea Able to Improve the Generalization of Deep Neural Networks on the Standard Supervised Learning?
- A neural network model of perception and reasoning
- RefBERT: Compressing BERT by Referencing to Pre-computed Representations
- Generic Bounds on the Maximum Deviations in Sequential Prediction: An Information-Theoretic Analysis
- Information Scaling Law of Deep Neural Networks
- Diagnosis and Analysis of Celiac Disease and Environmental Enteropathy on Biopsy Images using Deep Learning Approaches
- Verifiability and Predictability: Interpreting Utilities of Network Architectures for Point Cloud Processing
- Preserved Structure Across Vector Space Representations
- Whitening and second order optimization both make information in the dataset unusable during training, and can reduce or prevent generalization
- Filter Grafting for Deep Neural Networks: Reason, Method, and Cultivation
- Causal Representation Learning for Context-Aware Face Transfer
- Understanding Feature Selection and Feature Memorization in Recurrent Neural Networks
- A mathematical theory of imperfect communication: Energy efficiency considerations in multi-level coding
- Lattice Representation Learning
- Data-driven Regularization via Racecar Training for Generalizing Neural Networks
- Neural Network Activation Quantization with Bitwise Information Bottlenecks
- Cross-modal Image Retrieval with Deep Mutual Information Maximization
- Nested Learning For Multi-Granular Tasks
- Concepts, Properties and an Approach for Compositional Generalization
- Fundamental Limitations in Sequential Prediction and Recursive Algorithms: Bounds via an Entropic Analysis
- Model Reduction of Shallow CNN Model for Reliable Deployment of Information Extraction from Medical Reports
- An Information Bottleneck Problem with Rényi's Entropy
- How isotropic kernels perform on simple invariants
- Analysis of Information Flow Through U-Nets
- Algebraically-Informed Deep Networks (AIDN): A Deep Learning Approach to Represent Algebraic Structures
- Fundamental Limits of Prediction, Generalization, and Recursion: An Entropic-Innovations Perspective
- Interactive Re-Fitting as a Technique for Improving Word Embeddings
- Geometry Perspective Of Estimating Learning Capability Of Neural Networks
- BRIEF: Backward Reduction of CNNs with Information Flow Analysis