Gradient Episodic Memory for Continual Learning
arXiv:1706.08840
Abstract
One major obstacle towards AI is the poor ability of models to solve new problems quicker, and without forgetting previously acquired knowledge. To better understand this issue, we study the problem of continual learning, where the model observes, once and one by one, examples concerning a sequence of tasks. First, we propose a set of metrics to evaluate models learning over a continuum of data. These metrics characterize models not only by their test accuracy, but also in terms of their ability to transfer knowledge across tasks. Second, we propose a model for continual learning, called Gradient Episodic Memory (GEM) that alleviates forgetting, while allowing beneficial transfer of knowledge to previous tasks. Our experiments on variants of the MNIST and CIFAR-100 datasets demonstrate the strong performance of GEM when compared to the state-of-the-art.
Published at NIPS 2017
Cited by in corpus (65)
- Riemannian Walk for Incremental Learning: Understanding Forgetting and Intransigence
- Adaptive Aggregation Networks for Class-Incremental Learning
- Incremental Learning for Semantic Segmentation of Large-Scale Remote Sensing Data
- Deep Generative Dual Memory Network for Continual Learning
- Lifelong Generative Modeling
- Incremental Object Detection via Meta-Learning
- Learning to Learn without Forgetting by Maximizing Transfer and Minimizing Interference
- Incremental Classifier Learning with Generative Adversarial Networks
- Selfless Sequential Learning
- Inexact-ADMM Based Federated Meta-Learning for Fast and Continual Edge Learning
- Composable Planning with Attributes
- End-to-End Incremental Learning
- An Investigation of Replay-based Approaches for Continual Learning
- CoVIO: Online Continual Learning for Visual-Inertial Odometry
- SupportNet: solving catastrophic forgetting in class incremental learning with support data
- RATT: Recurrent Attention to Transient Tasks for Continual Image Captioning
- Direction Concentration Learning: Enhancing Congruency in Machine Learning
- Meta Continual Learning
- Lifelong Learning of Graph Neural Networks for Open-World Node Classification
- Flattening Sharpness for Dynamic Gradient Projection Memory Benefits Continual Learning
- Learning to Continuously Optimize Wireless Resource In Episodically Dynamic Environment
- Lifelong Adaptive Machine Learning for Sensor-based Human Activity Recognition Using Prototypical Networks
- Using Adapters to Overcome Catastrophic Forgetting in End-to-End Automatic Speech Recognition
- Continual learning of longitudinal health records
- AirLoop: Lifelong Loop Closure Detection
- CoDEPS: Online Continual Learning for Depth Estimation and Panoptic Segmentation
- ACE: Adapting to Changing Environments for Semantic Segmentation
- Distillation Techniques for Pseudo-rehearsal Based Incremental Learning
- Continual Learning for Monolingual End-to-End Automatic Speech Recognition
- Online Deep Learning based on Auto-Encoder
- Rethinking Continual Learning for Autonomous Agents and Robots
- Continual Learning with Neuromorphic Computing: Foundations, Methods, and Emerging Applications
- ARCADe: A Rapid Continual Anomaly Detector
- Weight Averaging: A Simple Yet Effective Method to Overcome Catastrophic Forgetting in Automatic Speech Recognition
- Online Distillation with Continual Learning for Cyclic Domain Shifts
- A Study of Continual Learning Methods for Q-Learning
- On the importance of cross-task features for class-incremental learning
- Representative Task Self-selection for Flexible Clustered Lifelong Learning
- Lifelong Language Knowledge Distillation
- BI-MAML: Balanced Incremental Approach for Meta Learning
- How neural networks find generalizable solutions: Self-tuned annealing in deep learning
- Deep Bilevel Learning
- Block Contextual MDPs for Continual Learning
- Federated Learning with Fair Averaging
- On Catastrophic Interference in Atari 2600 Games
- Regularize, Expand and Compress: Multi-task based Lifelong Learning via NonExpansive AutoML
- Lifelong Vehicle Trajectory Prediction Framework Based on Generative Replay
- Frosting Weights for Better Continual Training
- Overcoming Catastrophic Forgetting by Soft Parameter Pruning
- A Conceptual Framework for Lifelong Learning
- AlterSGD: Finding Flat Minima for Continual Learning by Alternative Training
- Continual Learning With Quasi-Newton Methods
- Group and Exclusive Sparse Regularization-based Continual Learning of CNNs
- Bilevel Continual Learning
- Variable-Shot Adaptation for Online Meta-Learning
- Dynamic VAEs with Generative Replay for Continual Zero-shot Learning
- ZS-IL: Looking Back on Learned Experiences For Zero-Shot Incremental Learning
- Weight Friction: A Simple Method to Overcome Catastrophic Forgetting and Enable Continual Learning
- Towards Making Deep Transfer Learning Never Hurt
- Split-and-Bridge: Adaptable Class Incremental Learning within a Single Neural Network
- Continual Active Learning for Efficient Adaptation of Machine Learning Models to Changing Image Acquisition
- Adversarial Incremental Learning
- Exploring the Challenges towards Lifelong Fact Learning
- Shared and Private VAEs with Generative Replay for Continual Learning
- MyMigrationBot: A Cloud-based Facebook Social Chatbot for Migrant Populations