ReduNet: A White-box Deep Network from the Principle of Maximizing Rate Reduction
arXiv:2105.10446
Abstract
This work attempts to provide a plausible theoretical framework that aims to interpret modern deep (convolutional) networks from the principles of data compression and discriminative representation. We argue that for high-dimensional multi-class data, the optimal linear discriminative representation maximizes the coding rate difference between the whole dataset and the average of all the subsets. We show that the basic iterative gradient ascent scheme for optimizing the rate reduction objective naturally leads to a multi-layer deep network, named ReduNet, which shares common characteristics of modern deep networks. The deep layered architectures, linear and nonlinear operators, and even parameters of the network are all explicitly constructed layer-by-layer via forward propagation, although they are amenable to fine-tuning via back propagation. All components of so-obtained "white-box" network have precise optimization, statistical, and geometric interpretation. Moreover, all linear operators of the so-derived network naturally become multi-channel convolutions when we enforce classification to be rigorously shift-invariant. The derivation in the invariant setting suggests a trade-off between sparsity and invariance, and also indicates that such a deep convolution network is significantly more efficient to construct and learn in the spectral domain. Our preliminary simulations and experiments clearly verify the effectiveness of both the rate reduction objective and the associated ReduNet. All code and data are available at \url{https://github.com/Ma-Lab-Berkeley}.
This paper integrates previous two manuscripts: arXiv:2006.08558 and arXiv:2010.14765, with significantly improved organization, presentation, and new results; V2 polishes writing and adds citation; V3 polishes writing, adds citation and experiments
References in corpus (16)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Neural Architecture Search with Reinforcement Learning
- Prevalence of Neural Collapse during the terminal phase of deep learning training
- Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
- Recovery Guarantees for One-hidden-layer Neural Networks
- Accelerated Gradient Descent Escapes Saddle Points Faster than Gradient Descent
- A Geometric Analysis of Neural Collapse with Unconstrained Features
- Interpretable Recurrent Neural Networks Using Sequential Sparse Recovery
- Are Pre-trained Convolutions Better than Pre-trained Transformers?
- Deep Sparse Subspace Clustering
- Deep Isometric Learning for Visual Recognition
- The Intrinsic Dimension of Images and Its Impact on Learning
- Dropout as a Low-Rank Regularizer for Matrix Factorization
- A Rate-Distortion Framework for Explaining Neural Network Decisions
- Convolutional Normalization: Improving Deep Convolutional Network Robustness and Training
Cited by in corpus (4)
- Learning High-Dimensional Parametric Maps via Reduced Basis Adaptive Residual Networks
- Towards Mitigating Dimensional Collapse of Representations in Collaborative Filtering
- Geometric Understanding of Discriminability and Transferability for Visual Domain Adaptation
- How Powerful is Graph Convolution for Recommendation?