Paying More Attention to Attention: Improving the Performance of Convolutional Neural Networks via Attention Transfer
arXiv:1612.03928
Abstract
Attention plays a critical role in human visual experience. Furthermore, it has recently been demonstrated that attention can also play an important role in the context of applying artificial neural networks to a variety of tasks from fields such as computer vision and NLP. In this work we show that, by properly defining attention for convolutional neural networks, we can actually use this type of information in order to significantly improve the performance of a student CNN network by forcing it to mimic the attention maps of a powerful teacher network. To that end, we propose several novel methods of transferring attention, showing consistent improvement across a variety of datasets and convolutional neural network architectures. Code and models for our experiments are available at https://github.com/szagoruyko/attention-transfer
Cited by in corpus (126)
- Align before Fuse: Vision and Language Representation Learning with Momentum Distillation
- Adaptive Multi-Teacher Multi-level Knowledge Distillation
- Efficient Medical Image Segmentation Based on Knowledge Distillation
- Vision Transformer with Attentive Pooling for Robust Facial Expression Recognition
- Neural Attention Distillation: Erasing Backdoor Triggers from Deep Neural Networks
- Enable Deep Learning on Mobile Devices: Methods, Systems, and Applications
- Learning Hierarchical Attention for Weakly-supervised Chest X-Ray Abnormality Localization and Diagnosis
- Torchreid: A Library for Deep Learning Person Re-Identification in Pytorch
- Double Similarity Distillation for Semantic Image Segmentation
- Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self Distillation
- Memorizing Complementation Network for Few-Shot Class-Incremental Learning
- On Representation Knowledge Distillation for Graph Neural Networks
- Knowledge Transfer via Distillation of Activation Boundaries Formed by Hidden Neurons
- Compacting Deep Neural Networks for Internet of Things: Methods and Applications
- Object DGCNN: 3D Object Detection using Dynamic Graphs
- Relational Knowledge Distillation
- A Survey on Symbolic Knowledge Distillation of Large Language Models
- Distilling Object Detectors with Task Adaptive Regularization
- Improved training of binary networks for human pose estimation and image recognition
- Channel Distillation: Channel-Wise Attention for Knowledge Distillation
- Distilling Object Detectors with Feature Richness
- Variational Information Distillation for Knowledge Transfer
- Bridging the Gap Between Patient-specific and Patient-independent Seizure Prediction via Knowledge Distillation
- Towards Optimal Structured CNN Pruning via Generative Adversarial Learning
- Regularizing Class-wise Predictions via Self-knowledge Distillation
- ShrinkTeaNet: Million-scale Lightweight Face Recognition via Shrinking Teacher-Student Networks
- Role of Data Augmentation Strategies in Knowledge Distillation for Wearable Sensor Data
- Radio2Text: Streaming Speech Recognition Using mmWave Radio Signals
- DBP: Discrimination Based Block-Level Pruning for Deep Model Acceleration
- Doing More with Less: Overcoming Data Scarcity for POI Recommendation via Cross-Region Transfer
- Symmetric Uncertainty-Aware Feature Transmission for Depth Super-Resolution
- DarkRank: Accelerating Deep Metric Learning via Cross Sample Similarities Transfer
- Residual Knowledge Distillation
- Label-guided Attention Distillation for Lane Segmentation
- Multi-Label Knowledge Distillation
- Knowledge Amalgamation for Object Detection with Transformers
- Group channel pruning and spatial attention distilling for object detection
- Justifying Diagnosis Decisions by Deep Neural Networks
- Switchable Precision Neural Networks
- SCAN: A Scalable Neural Networks Framework Towards Compact and Efficient Models
- Maximizing Discrimination Capability of Knowledge Distillation with Energy Function
- Search for Better Students to Learn Distilled Knowledge
- Multi-View Attention Transfer for Efficient Speech Enhancement
- MimicDet: Bridging the Gap Between One-Stage and Two-Stage Object Detection
- Tackling Catastrophic Forgetting and Background Shift in Continual Semantic Segmentation
- Distilling Knowledge via Knowledge Review
- CATFace: Cross-Attribute-Guided Transformer with Self-Attention Distillation for Low-Quality Face Recognition
- Balanced Knowledge Distillation for Long-tailed Learning
- PyNET-QxQ: An Efficient PyNET Variant for QxQ Bayer Pattern Demosaicing in CMOS Image Sensors
- Learning Scene Structure Guidance via Cross-Task Knowledge Transfer for Single Depth Super-Resolution
- Semi-Supervised Domain Generalizable Person Re-Identification
- Self-supervised Spatio-temporal Representation Learning for Videos by Predicting Motion and Appearance Statistics
- Inter-Region Affinity Distillation for Road Marking Segmentation
- Topological Persistence Guided Knowledge Distillation for Wearable Sensor Data
- Leveraging Angular Distributions for Improved Knowledge Distillation
- Learning Metrics from Teachers: Compact Networks for Image Embedding
- Towards Cross-modality Medical Image Segmentation with Online Mutual Knowledge Distillation
- Full-Cycle Energy Consumption Benchmark for Low-Carbon Computer Vision
- Rethinking Implicit Neural Representations for Vision Learners
- NeuGuard: Lightweight Neuron-Guided Defense against Membership Inference Attacks
- Distillation Guided Residual Learning for Binary Convolutional Neural Networks
- Attribution Preservation in Network Compression for Reliable Network Interpretation
- Spatial Likelihood Voting with Self-Knowledge Distillation for Weakly Supervised Object Detection
- RNAS-CL: Robust Neural Architecture Search by Cross-Layer Knowledge Distillation
- Computation-Efficient Knowledge Distillation via Uncertainty-Aware Mixup
- Restructuring, Pruning, and Adjustment of Deep Models for Parallel Distributed Inference
- Distilling Ensemble of Explanations for Weakly-Supervised Pre-Training of Image Segmentation Models
- Biphasic Learning of GANs for High-Resolution Image-to-Image Translation
- Learning with Privileged Information for Efficient Image Super-Resolution
- Embedding Transfer with Label Relaxation for Improved Metric Learning
- Knowledge Distillation Under Ideal Joint Classifier Assumption
- Knowledge Transfer via Dense Cross-Layer Mutual-Distillation
- Distilling Audio-Visual Knowledge by Compositional Contrastive Learning
- Bidirectional Knowledge Reconfiguration for Lightweight Point Cloud Analysis
- Training convolutional neural networks with cheap convolutions and online distillation
- LEA-Net: Layer-wise External Attention Network for Efficient Color Anomaly Detection
- Robustness and Diversity Seeking Data-Free Knowledge Distillation
- Parameter Efficient Deep Neural Networks with Bilinear Projections
- Separable Layers Enable Structured Efficient Linear Substitutions
- LNPT: Label-free Network Pruning and Training
- Knowledge Distillation in Document Retrieval
- Student Network Learning via Evolutionary Knowledge Distillation
- Text is Text, No Matter What: Unifying Text Recognition using Knowledge Distillation
- Oracle Teacher: Leveraging Target Information for Better Knowledge Distillation of CTC Models
- Dual-path CNN with Max Gated block for Text-Based Person Re-identification
- Learning from a Lightweight Teacher for Efficient Knowledge Distillation
- Knowledge Distillation Meets Self-Supervision
- Information Theoretic Representation Distillation
- TinyGAN: Distilling BigGAN for Conditional Image Generation
- BNAS v2: Learning Architectures for Binary Networks with Empirical Improvements
- Distill-2MD-MTL: Data Distillation based on Multi-Dataset Multi-Domain Multi-Task Frame Work to Solve Face Related Tasksks, Multi Task Learning, Semi-Supervised Learning
- Fixing the Teacher-Student Knowledge Discrepancy in Distillation
- Dynamic Slimmable Denoising Network
- Adaptive Distillation: Aggregating Knowledge from Multiple Paths for Efficient Distillation
- Measure Twice, Cut Once: Quantifying Bias and Fairness in Deep Neural Networks
- Pool of Experts: Realtime Querying Specialized Knowledge in Massive Neural Networks
- Enhancing Data-Free Adversarial Distillation with Activation Regularization and Virtual Interpolation
- Deep Ensemble Collaborative Learning by using Knowledge-transfer Graph for Fine-grained Object Classification
- Zero-shot Adversarial Quantization
- EEG aided boosting of single-lead ECG based sleep staging with Deep Knowledge Distillation
- Video Playback Rate Perception for Self-supervisedSpatio-Temporal Representation Learning
- Prime-Aware Adaptive Distillation
- LabelEnc: A New Intermediate Supervision Method for Object Detection
- Introspective Learning by Distilling Knowledge from Online Self-explanation
- Knowledge Representing: Efficient, Sparse Representation of Prior Knowledge for Knowledge Distillation
- Creating Lightweight Object Detectors with Model Compression for Deployment on Edge Devices
- Semi-Online Knowledge Distillation
- Arch-Net: Model Distillation for Architecture Agnostic Model Deployment
- GenURL: A General Framework for Unsupervised Representation Learning
- Piggyback GAN: Efficient Lifelong Learning for Image Conditioned Generation
- A Computer Vision Approach to Combat Lyme Disease
- Unsupervised Domain Adaptive Person Re-Identification via Human Learning Imitation
- Topology Distillation for Recommender System
- SERE: Exploring Feature Self-relation for Self-supervised Transformer
- Confidence Conditioned Knowledge Distillation
- AIP: Adversarial Iterative Pruning Based on Knowledge Transfer for Convolutional Neural Networks
- Multi-Scale Aligned Distillation for Low-Resolution Detection
- 3D attention mechanism for fine-grained classification of table tennis strokes using a Twin Spatio-Temporal Convolutional Neural Networks
- A One-step Pruning-recovery Framework for Acceleration of Convolutional Neural Networks
- Sequence-to-Sequence Learning via Attention Transfer for Incremental Speech Recognition
- Transferring Inter-Class Correlation
- Distilling Visual Priors from Self-Supervised Learning
- Multilingual AMR Parsing with Noisy Knowledge Distillation
- CHEER: Rich Model Helps Poor Model via Knowledge Infusion
- Diversified Mutual Learning for Deep Metric Learning
- A Survey on Green Deep Learning