Embedding Adaptation is Still Needed for Few-Shot Learning
arXiv:2104.07255
Abstract
Constructing new and more challenging tasksets is a fruitful methodology to analyse and understand few-shot classification methods. Unfortunately, existing approaches to building those tasksets are somewhat unsatisfactory: they either assume train and test task distributions to be identical -- which leads to overly optimistic evaluations -- or take a "worst-case" philosophy -- which typically requires additional human labor such as obtaining semantic class relationships. We propose ATG, a principled clustering method to defining train and test tasksets without additional human knowledge. ATG models train and test task distributions while requiring them to share a predefined amount of information. We empirically demonstrate the effectiveness of ATG in generating tasksets that are easier, in-between, or harder than existing benchmarks, including those that rely on semantic information. Finally, we leverage our generated tasksets to shed a new light on few-shot classification: gradient-based methods -- previously believed to underperform -- can outperform metric-based ones when transfer is most challenging.
In submission
References in corpus (8)
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Meta-SGD: Learning to Learn Quickly for Few-Shot Learning
- Efficient Off-Policy Meta-Reinforcement Learning via Probabilistic Context Variables
- A Large-scale Study of Representation Learning with the Visual Task Adaptation Benchmark
- Online Meta-Learning
- Understanding and Improving Information Transfer in Multi-Task Learning
- Meta-learning of Sequential Strategies
- The TCGA Meta-Dataset Clinical Benchmark