One to Transfer All: A Universal Transfer Framework for Vision Foundation Model with Few Data
arXiv:2111.12386
Abstract
The foundation model is not the last chapter of the model production pipeline. Transferring with few data in a general way to thousands of downstream tasks is becoming a trend of the foundation model's application. In this paper, we proposed a universal transfer framework: One to Transfer All (OTA) to transfer any Vision Foundation Model (VFM) to any downstream tasks with few downstream data. We first transfer a VFM to a task-specific model by Image Re-representation Fine-tuning (IRF) then distilling knowledge from a task-specific model to a deployed model with data produced by Downstream Image-Guided Generation (DIGG). OTA has no dependency on upstream data, VFM, and downstream tasks when transferring. It also provides a way for VFM researchers to release their upstream information for better transferring but not leaking data due to privacy requirements. Massive experiments validate the effectiveness and superiority of our methods in few data setting. Our code will be released.
Technical Report
References in corpus (12)
- Distilling the Knowledge in a Neural Network
- Rethinking Atrous Convolution for Semantic Image Segmentation
- Learning Transferable Visual Models From Natural Language Supervision
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- Language Models are Few-Shot Learners
- Learning Transferable Features with Deep Adaptation Networks
- On the Opportunities and Risks of Foundation Models
- BEiT: BERT Pre-Training of Image Transformers
- A Large-scale Study of Representation Learning with the Visual Task Adaptation Benchmark
- Self-supervised Knowledge Distillation for Few-shot Learning
- KD3A: Unsupervised Multi-Source Decentralized Domain Adaptation via Knowledge Distillation
- An Embarrassingly Simple Approach for Transfer Learning from Pretrained Language Models