High-Capacity Expert Binary Networks
arXiv:2010.03558
Abstract
Network binarization is a promising hardware-aware direction for creating efficient deep models. Despite its memory and computational advantages, reducing the accuracy gap between binary models and their real-valued counterparts remains an unsolved challenging research problem. To this end, we make the following 3 contributions: (a) To increase model capacity, we propose Expert Binary Convolution, which, for the first time, tailors conditional computing to binary networks by learning to select one data-specific expert binary filter at a time conditioned on input features. (b) To increase representation capacity, we propose to address the inherent information bottleneck in binary networks by introducing an efficient width expansion mechanism which keeps the binary operations within the same budget. (c) To improve network design, we propose a principled binary network growth mechanism that unveils a set of network topologies of favorable properties. Overall, our method improves upon prior work, with no increase in computational cost, by , reaching a groundbreaking on ImageNet classification. Code will be made available .
Accepted at ICLR 2021
References in corpus (14)
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- mixup: Beyond Empirical Risk Minimization
- Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
- The Concrete Distribution: A Continuous Relaxation of Discrete Random Variables
- Trained Ternary Quantization
- CondConv: Conditionally Parameterized Convolutions for Efficient Inference
- XNOR-Net++: Improved Binary Neural Networks
- Training Binary Neural Networks with Real-to-Binary Convolutions
- Latent Weights Do Not Exist: Rethinking Binarized Neural Network Optimization
- Learning to Branch for Multi-Task Learning
- Improved training of binary networks for human pose estimation and image recognition
- Checkmate: Breaking the Memory Wall with Optimal Tensor Rematerialization
- Accurate and Compact Convolutional Neural Networks with Trained Binarization
- Differentiable Architecture Search with Ensemble Gumbel-Softmax