Understanding Continual Learning Settings with Data Distribution Drift Analysis
arXiv:2104.01678
Abstract
Classical machine learning algorithms often assume that the data are drawn i.i.d. from a stationary probability distribution. Recently, continual learning emerged as a rapidly growing area of machine learning where this assumption is relaxed, i.e. where the data distribution is non-stationary and changes over time. This paper represents the state of data distribution by a context variable . A drift in leads to a data distribution drift. A context drift may change the target distribution, the input distribution, or both. Moreover, distribution drifts might be abrupt or gradual. In continual learning, context drifts may interfere with the learning process and erase previously learned knowledge; thus, continual learning algorithms must include specialized mechanisms to deal with such drifts. In this paper, we aim to identify and categorize different types of context drifts and potential assumptions about them, to better characterize various continual-learning scenarios. Moreover, we propose to use the distribution drift framework to provide more precise definitions of several terms commonly used in the continual learning field.
References in corpus (11)
- Distilling the Knowledge in a Neural Network
- PathNet: Evolution Channels Gradient Descent in Super Neural Networks
- Efficient Lifelong Learning with A-GEM
- Three scenarios for continual learning
- A Closer Look at Invalid Action Masking in Policy Gradient Algorithms
- Dive into Deep Learning
- Task Agnostic Continual Learning via Meta Learning
- DisCoRL: Continual Reinforcement Learning via Policy Distillation
- Optimal Continual Learning has Perfect Memory and is NP-hard
- Continuum: Simple Management of Complex Continual Learning Scenarios
- Continual Learning: Tackling Catastrophic Forgetting in Deep Neural Networks with Replay Processes