Multi-Scale Dense Networks for Resource Efficient Image Classification
arXiv:1703.09844
Abstract
In this paper we investigate image classification with computational resource limits at test time. Two such settings are: 1. anytime classification, where the network's prediction for a test example is progressively updated, facilitating the output of a prediction at any time; and 2. budgeted batch classification, where a fixed amount of computation is available to classify a set of examples that can be spent unevenly across "easier" and "harder" inputs. In contrast to most prior work, such as the popular Viola and Jones algorithm, our approach is based on convolutional neural networks. We train multiple classifiers with varying resource demands, which we adaptively apply during test time. To maximally re-use computation between the classifiers, we incorporate them as early-exits into a single deep convolutional neural network and inter-connect them with dense connectivity. To facilitate high quality classification early on, we use a two-dimensional multi-scale network architecture that maintains coarse and fine level features all-throughout the network. Experiments on three image-classification tasks demonstrate that our framework substantially improves the existing state-of-the-art in both settings.
Cited by in corpus (85)
- Split Computing and Early Exiting for Deep Learning Applications: Survey and Research Challenges
- SPINN: Synergistic Progressive Inference of Neural Networks over Device and Cloud
- Graph HyperNetworks for Neural Architecture Search
- Benchmarking Neural Network Robustness to Common Corruptions and Surface Variations
- Adaptive Inference through Early-Exit Networks: Design, Challenges and Directions
- Once-for-All: Train One Network and Specialize it for Efficient Deployment
- Dual Dynamic Inference: Enabling More Efficient, Adaptive and Controllable Deep Inference
- Big-Little Net: An Efficient Multi-Scale Feature Representation for Visual and Speech Recognition
- MobiSR: Efficient On-Device Super-Resolution through Heterogeneous Mobile Processors
- Dynamic Convolution: Attention over Convolution Kernels
- Learning Enriched Features for Real Image Restoration and Enhancement
- Multi-Scale Boosted Dehazing Network with Dense Feature Fusion
- Depth-Adaptive Transformer
- Pruning and Quantization for Deep Neural Network Acceleration: A Survey
- Practical Solutions for Machine Learning Safety in Autonomous Vehicles
- Deep convolutional neural networks for multi-planar lung nodule detection: improvement in small nodule identification
- Batch-Shaping for Learning Conditional Channel Gated Networks
- On-Device Machine Learning: An Algorithms and Learning Theory Perspective
- Dynamic Neural Networks: A Survey
- Triple Wins: Boosting Accuracy, Robustness and Efficiency Together by Enabling Input-Adaptive Inference
- Weight-Sharing Neural Architecture Search: A Battle to Shrink the Optimization Gap
- Revisiting Locally Supervised Learning: an Alternative to End-to-end Training
- Resolution Adaptive Networks for Efficient Inference
- It's always personal: Using Early Exits for Efficient On-Device CNN Personalisation
- Wisdom of Committees: An Overlooked Approach To Faster and More Accurate Models
- Once-for-All Adversarial Training: In-Situ Tradeoff between Robustness and Accuracy for Free
- Deep Learning Through the Lens of Example Difficulty
- MicroNet: Towards Image Recognition with Extremely Low FLOPs
- Not All Images are Worth 16x16 Words: Dynamic Transformers for Efficient Image Recognition
- Efficient Visual Recognition with Deep Neural Networks: A Survey on Recent Advances and New Directions
- A Panda? No, It's a Sloth: Slowdown Attacks on Adaptive Multi-Exit Neural Network Inference
- Universally Slimmable Networks and Improved Training Techniques
- ELF: An Early-Exiting Framework for Long-Tailed Classification
- Anytime Inference with Distilled Hierarchical Neural Ensembles
- AdaViT: Adaptive Vision Transformers for Efficient Image Recognition
- SCAN: A Scalable Neural Networks Framework Towards Compact and Efficient Models
- Learning Dynamic Routing for Semantic Segmentation
- Scaling-Translation-Equivariant Networks with Decomposed Convolutional Filters
- Activate or Not: Learning Customized Activation
- ApproxNet: Content and Contention-Aware Video Analytics System for Embedded Clients
- Hard-Attention for Scalable Image Classification
- Scalable Transformers for Neural Machine Translation
- S2DNAS:Transforming Static CNN Model for Dynamic Inference via Neural Architecture Search
- ApproxDet: Content and Contention-Aware Approximate Object Detection for Mobiles
- Fractional Skipping: Towards Finer-Grained Dynamic CNN Inference
- RomeBERT: Robust Training of Multi-Exit BERT
- Densely connected multidilated convolutional networks for dense prediction tasks
- Computation on Sparse Neural Networks: an Inspiration for Future Hardware
- 2D or not 2D? Adaptive 3D Convolution Selection for Efficient Video Recognition
- Students are the Best Teacher: Exit-Ensemble Distillation with Multi-Exits
- Fingerprinting Multi-exit Deep Neural Network Models via Inference Time
- Knowledge Transfer via Dense Cross-Layer Mutual-Distillation
- Dynamic Network Quantization for Efficient Video Inference
- Learning Sparse Mixture of Experts for Visual Question Answering
- MicroNet: Improving Image Recognition with Extremely Low FLOPs
- CoDiNet: Path Distribution Modeling with Consistency and Diversity for Dynamic Routing
- Dynamic Inference: A New Approach Toward Efficient Video Action Recognition
- Embedded Knowledge Distillation in Depth-Level Dynamic Neural Network
- DS-Net++: Dynamic Weight Slicing for Efficient Inference in CNNs and Transformers
- Improving Anytime Prediction with Parallel Cascaded Networks and a Temporal-Difference Loss
- Consumer Image Quality Prediction using Recurrent Neural Networks for Spatial Pooling
- Recurrent Convolution for Compact and Cost-Adjustable Neural Networks: An Empirical Study
- Dynamic Sparsity Neural Networks for Automatic Speech Recognition
- Dynamic Slimmable Denoising Network
- Dynamic Resolution Network
- Temporally Resolution Decrement: Utilizing the Shape Consistency for Higher Computational Efficiency
- ParaDiS: Parallelly Distributable Slimmable Neural Networks
- Rapid Elastic Architecture Search under Specialized Classes and Resource Constraints
- BasisNet: Two-stage Model Synthesis for Efficient Inference
- MOOD: Multi-level Out-of-distribution Detection
- Improved Techniques for Quantizing Deep Networks with Adaptive Bit-Widths
- Dynamic Slimmable Network
- Dynamic Routing Networks
- Back-Projection Pipeline
- Localizing Interpretable Multi-scale informative Patches Derived from Media Classification Task
- Learning to Generate Content-Aware Dynamic Detectors
- Dynamic Parameterized Network for CTR Prediction
- Joslim: Joint Widths and Weights Optimization for Slimmable Neural Networks
- Hierarchical Action Classification with Network Pruning
- Auto-Split: A General Framework of Collaborative Edge-Cloud AI
- A Survey on Green Deep Learning
- Consistent Accelerated Inference via Confident Adaptive Transformers
- Self-Regulation for Semantic Segmentation
- Prioritized Subnet Sampling for Resource-Adaptive Supernet Training
- Dynamic Domain Adaptation for Efficient Inference