CPT: Colorful Prompt Tuning for Pre-trained Vision-Language Models
arXiv:2109.11797
Abstract
Pre-Trained Vision-Language Models (VL-PTMs) have shown promising capabilities in grounding natural language in image data, facilitating a broad variety of cross-modal tasks. However, we note that there exists a significant gap between the objective forms of model pre-training and fine-tuning, resulting in a need for large amounts of labeled data to stimulate the visual grounding capability of VL-PTMs for downstream tasks. To address the challenge, we present Cross-modal Prompt Tuning (CPT, alternatively, Colorful Prompt Tuning), a novel paradigm for tuning VL-PTMs, which reformulates visual grounding into a fill-in-the-blank problem with color-based co-referential markers in image and text, maximally mitigating the gap. In this way, CPT enables strong few-shot and even zero-shot visual grounding capabilities of VL-PTMs. Comprehensive experimental results show that the prompt-tuned VL-PTMs outperform their fine-tuned counterparts by a large margin (e.g., 17.3% absolute accuracy improvement, and 73.8% relative standard deviation reduction on average with one shot in RefCOCO evaluation). We make the data and code for this paper publicly available at https://github.com/thunlp/CPT.
Work in progress
References in corpus (11)
- Learning Transferable Visual Models From Natural Language Supervision
- Learning to Prompt for Vision-Language Models
- Language Models are Few-Shot Learners
- Zero-Shot Text-to-Image Generation
- SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
- Learning How to Ask: Querying LMs with Mixtures of Soft Prompts
- Prefix-Tuning: Optimizing Continuous Prompts for Generation
- Fine-Tuning Pretrained Language Models: Weight Initializations, Data Orders, and Early Stopping
- Multimodal Few-Shot Learning with Frozen Language Models
- Calibrate Before Use: Improving Few-Shot Performance of Language Models
- Few-shot Object Grounding and Mapping for Natural Language Robot Instruction Following
Cited by in corpus (5)
- CLIP-Adapter: Better Vision-Language Models with Feature Adapters
- Dual Modality Prompt Tuning for Vision-Language Pre-Trained Model
- DenseCLIP: Language-Guided Dense Prediction with Context-Aware Prompting
- Fine-grained Visual-Text Prompt-Driven Self-Training for Open-Vocabulary Object Detection
- MuDPT: Multi-modal Deep-symphysis Prompt Tuning for Large Pre-trained Vision-Language Models