CLIP-Adapter: Better Vision-Language Models with Feature Adapters
arXiv:2110.04544
Abstract
Large-scale contrastive vision-language pre-training has shown significant progress in visual representation learning. Unlike traditional visual systems trained by a fixed set of discrete labels, a new paradigm was introduced in \cite{radford2021learning} to directly learn to align images with raw texts in an open-vocabulary setting. On downstream tasks, a carefully chosen text prompt is employed to make zero-shot predictions.~To avoid non-trivial prompt engineering, context optimization \cite{zhou2021coop} has been proposed to learn continuous vectors as task-specific prompts with few-shot training examples.~In this paper, we show that there is an alternative path to achieve better vision-language models other than prompt tuning.~While prompt tuning is for the textual inputs, we propose CLIP-Adapter to conduct fine-tuning with feature adapters on either visual or language branch. Specifically, CLIP-Adapter adopts an additional bottleneck layer to learn new features and performs residual-style feature blending with the original pre-trained features.~As a consequence, CLIP-Adapter is able to outperform context optimization while maintains a simple design. Experiments and extensive ablation studies on various visual classification tasks demonstrate the effectiveness of our approach. Code is released at t https://github.com/gaopengcuhk/CLIP-Adapter.
Accepted by IJCV
References in corpus (8)
- Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
- MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- Learning to Prompt for Vision-Language Models
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
- Fine-Grained Visual Classification of Aircraft
- CPT: Colorful Prompt Tuning for Pre-trained Vision-Language Models
- Dual-stream Network for Visual Recognition
Cited by in corpus (6)
- Tip-Adapter: Training-free CLIP-Adapter for Better Vision-Language Modeling
- Dual-stream Network for Visual Recognition
- CLAMP: Prompt-based Contrastive Learning for Connecting Language and Animal Pose
- Fine-grained Visual-Text Prompt-Driven Self-Training for Open-Vocabulary Object Detection
- Queryable Prototype Multiple Instance Learning with Vision-Language Models for Incremental Whole Slide Image Classification
- Uni-Perceiver: Pre-training Unified Architecture for Generic Perception for Zero-shot and Few-shot Tasks