TAG: Task-based Accumulated Gradients for Lifelong learning
arXiv:2105.05155
Abstract
When an agent encounters a continual stream of new tasks in the lifelong learning setting, it leverages the knowledge it gained from the earlier tasks to help learn the new tasks better. In such a scenario, identifying an efficient knowledge representation becomes a challenging problem. Most research works propose to either store a subset of examples from the past tasks in a replay buffer, dedicate a separate set of parameters to each task or penalize excessive updates over parameters by introducing a regularization term. While existing methods employ the general task-agnostic stochastic gradient descent update rule, we propose a task-aware optimizer that adapts the learning rate based on the relatedness among tasks. We utilize the directions taken by the parameters during the updates by accumulating the gradients specific to each task. These task-based accumulated gradients act as a knowledge base that is maintained and updated throughout the stream. We empirically show that our proposed adaptive learning rate not only accounts for catastrophic forgetting but also allows positive backward transfer. We also show that our method performs better than several state-of-the-art methods in lifelong learning on complex datasets with a large number of tasks.
Published at 1st Conference on Lifelong Learning Agents, 2022
References in corpus (18)
- Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithms
- A continual learning survey: Defying forgetting in classification tasks
- Weight Uncertainty in Neural Networks
- Efficient Lifelong Learning with A-GEM
- Three scenarios for continual learning
- Improving Generalization Performance by Switching from Adam to SGD
- On Tiny Episodic Memories in Continual Learning
- Re-evaluating Continual Learning Scenarios: A Categorization and Case for Strong Baselines
- Uncertainty-based Continual Learning with Adaptive Regularization
- Understanding the Role of Training Regimes in Continual Learning
- Continual Unsupervised Representation Learning
- Learn to Grow: A Continual Structure Learning Framework for Overcoming Catastrophic Forgetting
- Convergence Analysis of Proximal Gradient with Momentum for Nonconvex Optimization
- La-MAML: Look-ahead Meta Learning for Continual Learning
- Continual Learning with Adaptive Weights (CLAW)
- Towards Training Recurrent Neural Networks for Lifelong Learning
- Towards Understanding Generalization in Gradient-Based Meta-Learning
- Generalisation Guarantees for Continual Learning with Orthogonal Gradient Descent