Probing transfer learning with a model of synthetic correlated datasets
arXiv:2106.05418 · doi:10.1088/2632-2153/ac4f3f
Abstract
Transfer learning can significantly improve the sample efficiency of neural networks, by exploiting the relatedness between a data-scarce target task and a data-abundant source task. Despite years of successful applications, transfer learning practice often relies on ad-hoc solutions, while theoretical understanding of these procedures is still limited. In the present work, we re-think a solvable model of synthetic data as a framework for modeling correlation between data-sets. This setup allows for an analytic characterization of the generalization performance obtained when transferring the learned feature map from the source to the target task. Focusing on the problem of training two-layer networks in a binary classification setting, we show that our model can capture a range of salient features of transfer learning with real data. Moreover, by exploiting parametric control over the correlation between the two data-sets, we systematically investigate under which conditions the transfer of features is beneficial for generalization.
References in corpus (14)
- How transferable are features in deep neural networks?
- Language Models are Few-Shot Learners
- On the Opportunities and Risks of Foundation Models
- Theoretical Models of Learning to Learn
- A Model of Inductive Bias Learning
- What is being transferred in transfer learning?
- Geometric Dataset Distances via Optimal Transport
- Double Trouble in Double Descent : Bias and Variance(s) in the Lazy Regime
- Towards Automated Melanoma Screening: Exploring Transfer Learning Schemes
- Universality Laws for High-Dimensional Learning with Random Features
- DeepRadiologyNet: Radiologist Level Pathology Detection in CT Head Images
- Phase Transitions in Transfer Learning for High-Dimensional Perceptrons
- Similarity of Classification Tasks
- An Information-Geometric Distance on the Space of Tasks
Cited by in corpus (6)
- Gaussian Universality of Perceptrons with Random Labels
- A Farewell to the Bias-Variance Tradeoff? An Overview of the Theory of Overparameterized Machine Learning
- Daydreaming Hopfield Networks and their surprising effectiveness on correlated data
- The impact of memory on learning sequence-to-sequence tasks
- Optimal Protocols for Continual Learning via Statistical Physics and Control Theory
- The RL Perceptron: Generalisation Dynamics of Policy Learning in High Dimensions