Probing transfer learning with a model of synthetic correlated datasets
arXiv:2106.05418 · doi:10.1088/2632-2153/ac4f3f
Abstract
Transfer learning can significantly improve the sample efficiency of neural networks, by exploiting the relatedness between a data-scarce target task and a data-abundant source task. Despite years of successful applications, transfer learning practice often relies on ad-hoc solutions, while theoretical understanding of these procedures is still limited. In the present work, we re-think a solvable model of synthetic data as a framework for modeling correlation between data-sets. This setup allows for an analytic characterization of the generalization performance obtained when transferring the learned feature map from the source to the target task. Focusing on the problem of training two-layer networks in a binary classification setting, we show that our model can capture a range of salient features of transfer learning with real data. Moreover, by exploiting parametric control over the correlation between the two data-sets, we systematically investigate under which conditions the transfer of features is beneficial for generalization.
References in corpus (20)
- How transferable are features in deep neural networks?
- Language Models are Few-Shot Learners
- On the Opportunities and Risks of Foundation Models
- Theoretical Models of Learning to Learn
- A Model of Inductive Bias Learning
- What is being transferred in transfer learning?
- Generalisation error in learning with random features and the hidden manifold model
- Geometric Dataset Distances via Optimal Transport
- On the Theory of Transfer Learning: The Importance of Task Diversity
- Few-Shot Learning via Learning the Representation, Provably
- Double Trouble in Double Descent : Bias and Variance(s) in the Lazy Regime
- The Gaussian equivalence of generative models for learning with shallow neural networks
- Towards Automated Melanoma Screening: Exploring Transfer Learning Schemes
- Statistical learning theory of structured data
- Universality Laws for High-Dimensional Learning with Random Features
- DeepRadiologyNet: Radiologist Level Pathology Detection in CT Head Images
- Phase Transitions in Transfer Learning for High-Dimensional Perceptrons
- Similarity of Classification Tasks
- An Information-Geometric Distance on the Space of Tasks
- Double Double Descent: On Generalization Errors in Transfer Learning between Linear Regression Tasks
Cited by in corpus (6)
- Gaussian Universality of Perceptrons with Random Labels
- A Farewell to the Bias-Variance Tradeoff? An Overview of the Theory of Overparameterized Machine Learning
- Daydreaming Hopfield Networks and their surprising effectiveness on correlated data
- The impact of memory on learning sequence-to-sequence tasks
- Optimal Protocols for Continual Learning via Statistical Physics and Control Theory
- The RL Perceptron: Generalisation Dynamics of Policy Learning in High Dimensions