Sparse Flows: Pruning Continuous-depth Models
arXiv:2106.12718
Abstract
Continuous deep learning architectures enable learning of flexible probabilistic models for predictive modeling as neural ordinary differential equations (ODEs), and for generative modeling as continuous normalizing flows. In this work, we design a framework to decipher the internal dynamics of these continuous depth models by pruning their network architectures. Our empirical results suggest that pruning improves generalization for neural ODEs in generative modeling. We empirically show that the improvement is because pruning helps avoid mode-collapse and flatten the loss surface. Moreover, pruning finds efficient neural ODE representations with up to 98% less parameters compared to the original network, without loss of accuracy. We hope our results will invigorate further research into the performance-size trade-offs of modern continuous-depth models.
NeurIPS 2021
References in corpus (14)
- Distilling the Knowledge in a Neural Network
- NICE: Non-linear Independent Components Estimation
- Pruning Filters for Efficient ConvNets
- MADE: Masked Autoencoder for Distribution Estimation
- An Investigation into Neural Net Optimization via Hessian Eigenvalue Density
- Dissecting Neural ODEs
- Lipschitz Recurrent Neural Networks
- Coupling-based Invertible Neural Networks Are Universal Diffeomorphism Approximators
- Sum-of-Squares Polynomial Flow
- Causal Navigation by Continuous-time Neural Networks
- TorchDyn: A Neural Differential Equations Library
- The Expressive Power of a Class of Normalizing Flow Models
- Convex Potential Flows: Universal Probability Distributions with Optimal Transport and Convex Optimization
- Compressing Neural Networks: Towards Determining the Optimal Layer-wise Decomposition