Paying More Attention to Attention: Improving the Performance of Convolutional Neural Networks via Attention Transfer
arXiv:1612.03928
Abstract
Attention plays a critical role in human visual experience. Furthermore, it has recently been demonstrated that attention can also play an important role in the context of applying artificial neural networks to a variety of tasks from fields such as computer vision and NLP. In this work we show that, by properly defining attention for convolutional neural networks, we can actually use this type of information in order to significantly improve the performance of a student CNN network by forcing it to mimic the attention maps of a powerful teacher network. To that end, we propose several novel methods of transferring attention, showing consistent improvement across a variety of datasets and convolutional neural network architectures. Code and models for our experiments are available at https://github.com/szagoruyko/attention-transfer
Cited by in corpus (96)
- Align before Fuse: Vision and Language Representation Learning with Momentum Distillation
- Adaptive Multi-Teacher Multi-level Knowledge Distillation
- Efficient Medical Image Segmentation Based on Knowledge Distillation
- Neural Attention Distillation: Erasing Backdoor Triggers from Deep Neural Networks
- Enable Deep Learning on Mobile Devices: Methods, Systems, and Applications
- Learning Hierarchical Attention for Weakly-supervised Chest X-Ray Abnormality Localization and Diagnosis
- Torchreid: A Library for Deep Learning Person Re-Identification in Pytorch
- Double Similarity Distillation for Semantic Image Segmentation
- Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self Distillation
- Knowledge Transfer via Distillation of Activation Boundaries Formed by Hidden Neurons
- Compacting Deep Neural Networks for Internet of Things: Methods and Applications
- Object DGCNN: 3D Object Detection using Dynamic Graphs
- Relational Knowledge Distillation
- Distilling Object Detectors with Task Adaptive Regularization
- Improved training of binary networks for human pose estimation and image recognition
- Channel Distillation: Channel-Wise Attention for Knowledge Distillation
- Distilling Object Detectors with Feature Richness
- Variational Information Distillation for Knowledge Transfer
- Towards Optimal Structured CNN Pruning via Generative Adversarial Learning
- Regularizing Class-wise Predictions via Self-knowledge Distillation
- ShrinkTeaNet: Million-scale Lightweight Face Recognition via Shrinking Teacher-Student Networks
- Role of Data Augmentation Strategies in Knowledge Distillation for Wearable Sensor Data
- DBP: Discrimination Based Block-Level Pruning for Deep Model Acceleration
- DarkRank: Accelerating Deep Metric Learning via Cross Sample Similarities Transfer
- Residual Knowledge Distillation
- Switchable Precision Neural Networks
- Justifying Diagnosis Decisions by Deep Neural Networks
- SCAN: A Scalable Neural Networks Framework Towards Compact and Efficient Models
- Search for Better Students to Learn Distilled Knowledge
- MimicDet: Bridging the Gap Between One-Stage and Two-Stage Object Detection
- Multi-View Attention Transfer for Efficient Speech Enhancement
- Distilling Knowledge via Knowledge Review
- Tackling Catastrophic Forgetting and Background Shift in Continual Semantic Segmentation
- Balanced Knowledge Distillation for Long-tailed Learning
- Semi-Supervised Domain Generalizable Person Re-Identification
- Self-supervised Spatio-temporal Representation Learning for Videos by Predicting Motion and Appearance Statistics
- Inter-Region Affinity Distillation for Road Marking Segmentation
- Learning Scene Structure Guidance via Cross-Task Knowledge Transfer for Single Depth Super-Resolution
- Towards Cross-modality Medical Image Segmentation with Online Mutual Knowledge Distillation
- Learning Metrics from Teachers: Compact Networks for Image Embedding
- Distillation Guided Residual Learning for Binary Convolutional Neural Networks
- Full-Cycle Energy Consumption Benchmark for Low-Carbon Computer Vision
- Spatial Likelihood Voting with Self-Knowledge Distillation for Weakly Supervised Object Detection
- Computation-Efficient Knowledge Distillation via Uncertainty-Aware Mixup
- Attribution Preservation in Network Compression for Reliable Network Interpretation
- Restructuring, Pruning, and Adjustment of Deep Models for Parallel Distributed Inference
- Learning with Privileged Information for Efficient Image Super-Resolution
- Embedding Transfer with Label Relaxation for Improved Metric Learning
- Distilling Ensemble of Explanations for Weakly-Supervised Pre-Training of Image Segmentation Models
- Biphasic Learning of GANs for High-Resolution Image-to-Image Translation
- Knowledge Transfer via Dense Cross-Layer Mutual-Distillation
- Distilling Audio-Visual Knowledge by Compositional Contrastive Learning
- Student Network Learning via Evolutionary Knowledge Distillation
- Training convolutional neural networks with cheap convolutions and online distillation
- Knowledge Distillation in Document Retrieval
- LEA-Net: Layer-wise External Attention Network for Efficient Color Anomaly Detection
- Parameter Efficient Deep Neural Networks with Bilinear Projections
- Robustness and Diversity Seeking Data-Free Knowledge Distillation
- Separable Layers Enable Structured Efficient Linear Substitutions
- BNAS v2: Learning Architectures for Binary Networks with Empirical Improvements
- Distill-2MD-MTL: Data Distillation based on Multi-Dataset Multi-Domain Multi-Task Frame Work to Solve Face Related Tasksks, Multi Task Learning, Semi-Supervised Learning
- TinyGAN: Distilling BigGAN for Conditional Image Generation
- Dual-path CNN with Max Gated block for Text-Based Person Re-identification
- Knowledge Distillation Meets Self-Supervision
- Learning from a Lightweight Teacher for Efficient Knowledge Distillation
- Fixing the Teacher-Student Knowledge Discrepancy in Distillation
- Text is Text, No Matter What: Unifying Text Recognition using Knowledge Distillation
- Prime-Aware Adaptive Distillation
- Adaptive Distillation: Aggregating Knowledge from Multiple Paths for Efficient Distillation
- Pool of Experts: Realtime Querying Specialized Knowledge in Massive Neural Networks
- Dynamic Slimmable Denoising Network
- LabelEnc: A New Intermediate Supervision Method for Object Detection
- Deep Ensemble Collaborative Learning by using Knowledge-transfer Graph for Fine-grained Object Classification
- Enhancing Data-Free Adversarial Distillation with Activation Regularization and Virtual Interpolation
- Introspective Learning by Distilling Knowledge from Online Self-explanation
- Video Playback Rate Perception for Self-supervisedSpatio-Temporal Representation Learning
- Knowledge Representing: Efficient, Sparse Representation of Prior Knowledge for Knowledge Distillation
- Zero-shot Adversarial Quantization
- EEG aided boosting of single-lead ECG based sleep staging with Deep Knowledge Distillation
- Creating Lightweight Object Detectors with Model Compression for Deployment on Edge Devices
- Measure Twice, Cut Once: Quantifying Bias and Fairness in Deep Neural Networks
- 3D attention mechanism for fine-grained classification of table tennis strokes using a Twin Spatio-Temporal Convolutional Neural Networks
- Sequence-to-Sequence Learning via Attention Transfer for Incremental Speech Recognition
- Multilingual AMR Parsing with Noisy Knowledge Distillation
- Multi-Scale Aligned Distillation for Low-Resolution Detection
- Piggyback GAN: Efficient Lifelong Learning for Image Conditioned Generation
- CHEER: Rich Model Helps Poor Model via Knowledge Infusion
- AIP: Adversarial Iterative Pruning Based on Knowledge Transfer for Convolutional Neural Networks
- Distilling Visual Priors from Self-Supervised Learning
- A One-step Pruning-recovery Framework for Acceleration of Convolutional Neural Networks
- Confidence Conditioned Knowledge Distillation
- Transferring Inter-Class Correlation
- Topology Distillation for Recommender System
- Diversified Mutual Learning for Deep Metric Learning
- A Computer Vision Approach to Combat Lyme Disease
- Unsupervised Domain Adaptive Person Re-Identification via Human Learning Imitation