M2KD: Multi-model and Multi-level Knowledge Distillation for Incremental Learning
arXiv:1904.01769
Abstract
Incremental learning targets at achieving good performance on new categories without forgetting old ones. Knowledge distillation has been shown critical in preserving the performance on old classes. Conventional methods, however, sequentially distill knowledge only from the last model, leading to performance degradation on the old classes in later incremental learning steps. In this paper, we propose a multi-model and multi-level knowledge distillation strategy. Instead of sequentially distilling knowledge only from the last model, we directly leverage all previous model snapshots. In addition, we incorporate an auxiliary distillation to further preserve knowledge encoded at the intermediate feature levels. To make the model more memory efficient, we adapt mask based pruning to reconstruct all previous models with a small memory footprint. Experiments on standard incremental learning benchmarks show that our method preserves the knowledge on old classes better and improves the overall performance over standard distillation techniques.
References in corpus (8)
- Distilling the Knowledge in a Neural Network
- Riemannian Walk for Incremental Learning: Understanding Forgetting and Intransigence
- Efficient Lifelong Learning with A-GEM
- Learning Efficient Convolutional Networks through Network Slimming
- DSD: Dense-Sparse-Dense Training for Deep Neural Networks
- Piggyback: Adapting a Single Network to Multiple Tasks by Learning to Mask Weights
- SupportNet: solving catastrophic forgetting in class incremental learning with support data
- Continual State Representation Learning for Reinforcement Learning using Generative Replay
Cited by in corpus (10)
- Knowledge Distillation: A Survey
- A Comprehensive Study of Class Incremental Learning Algorithms for Visual Tasks
- Maintaining Discrimination and Fairness in Class Incremental Learning
- Inherit with Distillation and Evolve with Contrast: Exploring Class Incremental Semantic Segmentation Without Exemplar Memory
- Online Continual Learning via the Knowledge Invariant and Spread-out Properties
- Tackling Catastrophic Forgetting and Background Shift in Continual Semantic Segmentation
- Faster ILOD: Incremental Learning for Object Detectors based on Faster RCNN
- Initial Classifier Weights Replay for Memoryless Class Incremental Learning
- Energy Aligning for Biased Models
- DIODE: Dilatable Incremental Object Detection