Faster gaze prediction with dense networks and Fisher pruning
arXiv:1801.05787
Abstract
Predicting human fixations from images has recently seen large improvements by leveraging deep representations which were pretrained for object recognition. However, as we show in this paper, these networks are highly overparameterized for the task of fixation prediction. We first present a simple yet principled greedy pruning method which we call Fisher pruning. Through a combination of knowledge distillation and Fisher pruning, we obtain much more runtime-efficient architectures for saliency prediction, achieving a 10x speedup for the same AUC performance as a state of the art network on the CAT2000 dataset. Speeding up single-image gaze prediction is important for many real-world applications, but it is also a crucial step in the development of video saliency models, where the amount of data to be processed is substantially larger.
References in corpus (1)
Cited by in corpus (50)
- The State of Sparsity in Deep Neural Networks
- Ablation Studies in Artificial Neural Networks
- TSViz: Demystification of Deep Learning Models for Time-Series Analysis
- What Do Compressed Deep Neural Networks Forget?
- Learning Sparse Networks Using Targeted Dropout
- Zero-Cost Proxies for Lightweight NAS
- Inspect, Understand, Overcome: A Survey of Practical Methods for AI Safety
- BRECQ: Pushing the Limit of Post-Training Quantization by Block Reconstruction
- Layer-compensated Pruning for Resource-constrained Convolutional Neural Networks
- Evaluating Efficient Performance Estimators of Neural Architectures
- Feature Map Transform Coding for Energy-Efficient CNN Inference
- Saliency Prediction in the Deep Learning Era: Successes, Limitations, and Future Challenges
- A Closer Look at Structured Pruning for Neural Network Compression
- A Gradient Flow Framework For Analyzing Network Pruning
- How Powerful are Performance Predictors in Neural Architecture Search?
- Single Shot Structured Pruning Before Training
- BlockSwap: Fisher-guided Block Substitution for Network Compression on a Budget
- Neural networks on microcontrollers: saving memory at inference via operator reordering
- On Network Science and Mutual Information for Explaining Deep Neural Networks
- Taxonomy of Saliency Metrics for Channel Pruning
- GAN Slimming: All-in-One GAN Compression by A Unified Optimization Framework
- Characterising Across-Stack Optimisations for Deep Convolutional Neural Networks
- Composition of Saliency Metrics for Channel Pruning with a Myopic Oracle
- Non-Differentiable Supervised Learning with Evolution Strategies and Hybrid Methods
- GradSign: Model Performance Inference with Theoretical Insights
- Learning Sparse Neural Networks via Sensitivity-Driven Regularization
- Artificial neural networks condensation: A strategy to facilitate adaption of machine learning in medical settings by reducing computational burden
- Differentiable Joint Pruning and Quantization for Hardware Efficiency
- Dynamic Neural Network Channel Execution for Efficient Training
- Dynamical Isometry: The Missing Ingredient for Neural Network Pruning
- Distilling with Performance Enhanced Students
- Training Neural Networks with Fixed Sparse Masks
- DECORE: Deep Compression with Reinforcement Learning
- ProxyBO: Accelerating Neural Architecture Search via Bayesian Optimization with Zero-cost Proxies
- SOSP: Efficiently Capturing Global Correlations by Second-Order Structured Pruning
- ESPN: Extremely Sparse Pruned Networks
- Model Size Reduction Using Frequency Based Double Hashing for Recommender Systems
- Implicit Filter Sparsification In Convolutional Neural Networks
- Temporal Saliency Adaptation in Egocentric Videos
- Convolutional neural networks compression with low rank and sparse tensor decompositions
- One Weight Bitwidth to Rule Them All
- Channel-wise pruning of neural networks with tapering resource constraint
- Dirichlet Pruning for Neural Network Compression
- NeuralScale: Efficient Scaling of Neurons for Resource-Constrained Deep Neural Networks
- TOCO: A Framework for Compressing Neural Network Models Based on Tolerance Analysis
- Zero-Cost Operation Scoring in Differentiable Architecture Search
- Parallel Blockwise Knowledge Distillation for Deep Neural Network Compression
- OrthoReg: Robust Network Pruning Using Orthonormality Regularization
- DARTS-PRIME: Regularization and Scheduling Improve Constrained Optimization in Differentiable NAS
- MixMix: All You Need for Data-Free Compression Are Feature and Data Mixing