Learning to Prompt for Vision-Language Models
arXiv:2109.01134 · doi:10.1007/s11263-022-01653-1
Abstract
Large pre-trained vision-language models like CLIP have shown great potential in learning representations that are transferable across a wide range of downstream tasks. Different from the traditional representation learning that is based mostly on discretized labels, vision-language pre-training aligns images and texts in a common feature space, which allows zero-shot transfer to a downstream task via prompting, i.e., classification weights are synthesized from natural language describing classes of interest. In this work, we show that a major challenge for deploying such models in practice is prompt engineering, which requires domain expertise and is extremely time-consuming -- one needs to spend a significant amount of time on words tuning since a slight change in wording could have a huge impact on performance. Inspired by recent advances in prompt learning research in natural language processing (NLP), we propose Context Optimization (CoOp), a simple approach specifically for adapting CLIP-like vision-language models for downstream image recognition. Concretely, CoOp models a prompt's context words with learnable vectors while the entire pre-trained parameters are kept fixed. To handle different image recognition tasks, we provide two implementations of CoOp: unified context and class-specific context. Through extensive experiments on 11 datasets, we demonstrate that CoOp requires as few as one or two shots to beat hand-crafted prompts with a decent margin and is able to gain significant improvements over prompt engineering with more shots, e.g., with 16 shots the average gain is around 15% (with the highest reaching over 45%). Despite being a learning-based approach, CoOp achieves superb domain generalization performance compared with the zero-shot model using hand-crafted prompts.
International Journal of Computer Vision (IJCV), 2022. Update: Adds results on the DOSCO (DOmain Shift in COntext) benchmark
References in corpus (14)
- Learning Transferable Visual Models From Natural Language Supervision
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- Learning to Prompt for Vision-Language Models
- On the Opportunities and Risks of Foundation Models
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
- Domain Generalization: A Survey
- Zero-Shot Learning Through Cross-Modal Transfer
- Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing
- Do ImageNet Classifiers Generalize to ImageNet?
- Florence: A New Foundation Model for Computer Vision
- Prefix-Tuning: Optimizing Continuous Prompts for Generation
- Measuring Robustness to Natural Distribution Shifts in Image Classification
- Self-supervised learning of visual features through embedding images into text topic spaces
- On-Device Domain Generalization
Cited by in corpus (104)
- Learning to Prompt for Vision-Language Models
- Domain Generalization: A Survey
- Unleashing the potential of prompt engineering for large language models
- CLIP-Driven Universal Model for Organ Segmentation and Tumor Detection
- Open-Vocabulary DETR with Conditional Matching
- Generative AI Literacy: Twelve Defining Competencies
- A Survey on Mixture of Experts in Large Language Models
- Graphologue: Exploring Large Language Model Responses with Interactive Diagrams
- Tip-Adapter: Training-free CLIP-Adapter for Better Vision-Language Modeling
- AdaCLIP: Adapting CLIP with Hybrid Learnable Prompts for Zero-Shot Anomaly Detection
- Getting pwn'd by AI: Penetration Testing with Large Language Models
- RePrompt: Automatic Prompt Editing to Refine AI-Generative Art Towards Precise Expressions
- CLIP-Adapter: Better Vision-Language Models with Feature Adapters
- RS5M and GeoRSCLIP: A Large Scale Vision-Language Dataset and A Large Vision-Language Model for Remote Sensing
- CPT: Colorful Prompt Tuning for Pre-trained Vision-Language Models
- Large Language Models for Next Point-of-Interest Recommendation
- Dual Modality Prompt Tuning for Vision-Language Pre-Trained Model
- CLIP in Medical Imaging: A Survey
- CLIP-Count: Towards Text-Guided Zero-Shot Object Counting
- UAVs Meet LLMs: Overviews and Perspectives Toward Agentic Low-Altitude Mobility
- Augmenting Low-Resource Text Classification with Graph-Grounded Pre-training and Prompting
- Open-Vocabulary SAM: Segment and Recognize Twenty-thousand Classes Interactively
- FVP: Fourier Visual Prompting for Source-Free Unsupervised Domain Adaptation of Medical Image Segmentation
- Vision-Language Models for Edge Networks: A Comprehensive Survey
- CLIP4STR: A Simple Baseline for Scene Text Recognition with Pre-trained Vision-Language Model
- Semi-Supervised Semantic Segmentation Based on Pseudo-Labels: A Survey
- Towards Explainable In-the-Wild Video Quality Assessment: A Database and a Language-Prompted Approach
- Learning without Forgetting for Vision-Language Models
- Visual Tuning
- Multimodal Parameter-Efficient Few-Shot Class Incremental Learning
- Unveiling the Underwater World: CLIP Perception Model-Guided Underwater Image Enhancement
- How Does Fine-Tuning Impact Out-of-Distribution Detection for Vision-Language Models?
- CLOOB: Modern Hopfield Networks with InfoLOOB Outperform CLIP
- CLAMP: Prompt-based Contrastive Learning for Connecting Language and Animal Pose
- When Geoscience Meets Foundation Models: Towards General Geoscience Artificial Intelligence System
- Towards Visual-Prompt Temporal Answering Grounding in Medical Instructional Video
- Fine-grained Visual-Text Prompt-Driven Self-Training for Open-Vocabulary Object Detection
- Improving deep learning with prior knowledge and cognitive models: A survey on enhancing explainability, adversarial robustness and zero-shot learning
- Beyond Traditional Teaching: The Potential of Large Language Models and Chatbots in Graduate Engineering Education
- Cross-Modal Adapter for Vision-Language Retrieval
- Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives
- Embedded Visual Prompt Tuning
- Multi-modal Attribute Prompting for Vision-Language Models
- Actor-agnostic Multi-label Action Recognition with Multi-modal Query
- Unsupervised Domain Adaption Harnessing Vision-Language Pre-training
- Prompting Disentangled Embeddings for Knowledge Graph Completion with Pre-trained Language Model
- Dual-Modal Attention-Enhanced Text-Video Retrieval with Triplet Partial Margin Contrastive Learning
- Respond to Change with Constancy: Instruction-tuning with LLM for Non-I.I.D. Network Traffic Classification
- Foundation Models and Transformers for Anomaly Detection: A Survey
- Language-Inspired Relation Transfer for Few-shot Class-Incremental Learning
- Personalized Federated Continual Learning via Multi-granularity Prompt
- CLIPCleaner: Cleaning Noisy Labels with CLIP
- Unseen No More: Unlocking the Potential of CLIP for Generative Zero-shot HOI Detection
- Predicting Class Distribution Shift for Reliable Domain Adaptive Object Detection
- SenCLIP: Enhancing zero-shot land-use mapping for Sentinel-2 with ground-level prompting
- Prompt Tuning with Soft Context Sharing for Vision-Language Models
- PIG: Prompt Images Guidance for Night-Time Scene Parsing
- Text-Region Matching for Multi-Label Image Recognition with Missing Labels
- Prompting Video-Language Foundation Models with Domain-specific Fine-grained Heuristics for Video Question Answering
- A User-Friendly Framework for Generating Model-Preferred Prompts in Text-to-Image Synthesis
- Uni-Perceiver: Pre-training Unified Architecture for Generic Perception for Zero-shot and Few-shot Tasks
- Class Balance Matters to Active Class-Incremental Learning
- GenCLIP: Generalizing CLIP Prompts for Zero-shot Anomaly Detection
- Towards Foundation Models and Few-Shot Parameter-Efficient Fine-Tuning for Volumetric Organ Segmentation
- MoP-CLIP: A Mixture of Prompt-Tuned CLIP Models for Domain Incremental Learning
- SIA-OVD: Shape-Invariant Adapter for Bridging the Image-Region Gap in Open-Vocabulary Detection
- ViTextVQA: A Large-Scale Visual Question Answering Dataset and a Novel Multimodal Feature Fusion Method for Vietnamese Text Comprehension in Images
- SpecXNet: A Dual-Domain Convolutional Network for Robust Deepfake Detection
- ReCLIP++: Learn to Rectify the Bias of CLIP for Unsupervised Semantic Segmentation
- AttriPrompter: Auto-Prompting with Attribute Semantics for Zero-shot Nuclei Detection via Visual-Language Pre-trained Models
- MuDPT: Multi-modal Deep-symphysis Prompt Tuning for Large Pre-trained Vision-Language Models
- Text to Image for Multi-Label Image Recognition with Joint Prompt-Adapter Learning
- RSRefSeg 2: Decoupling Referring Remote Sensing Image Segmentation with Foundation Models
- Text-Derived Relational Graph-Enhanced Network for Skeleton-Based Action Segmentation
- CapS-Adapter: Caption-based MultiModal Adapter in Zero-Shot Classification
- LPN: Language-guided Prototypical Network for few-shot classification
- Compositional Kronecker Context Optimization for Vision-Language Models
- Advancing Prompt Learning through an External Layer
- Controllable Navigation Instruction Generation with Chain of Thought Prompting
- SPARK: Self-supervised Personalized Real-time Monocular Face Capture
- HCVP: Leveraging Hierarchical Contrastive Visual Prompt for Domain Generalization
- KNN Transformer with Pyramid Prompts for Few-Shot Learning
- PartEdit: Fine-Grained Image Editing using Pre-Trained Diffusion Models
- The Unreasonable Effectiveness of Large Language-Vision Models for Source-free Video Domain Adaptation
- Generalized Contrastive Learning for Multi-Modal Retrieval and Ranking
- COLA: Context-aware Language-driven Test-time Adaptation
- Embracing Diversity: Interpretable Zero-shot classification beyond one vector per class
- Key-Value Pair-Free Continual Learner via Task-Specific Prompt-Prototype
- Prompt Disentanglement via Language Guidance and Representation Alignment for Domain Generalization
- Domain-generalizable Face Anti-Spoofing with Patch-based Multi-tasking and Artifact Pattern Conversion
- Image-to-Text Translation for Interactive Image Recognition: A Comparative User Study with Non-Expert Users
- VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations
- Contrastive Regularization over LoRA for Multimodal Biomedical Image Incremental Learning
- Semantically Guided Dynamic Visual Prototype Refinement for Compositional Zero-Shot Learning
- Visual Objectification in Films: Towards a New AI Task for Video Interpretation
- Localized Conformal Prediction for Image Classification with Vision-Language Models
- Unsupervised Data-Efficient Cross-Modal Retrieval with Global-Neighborhood Alignment Hashing
- Holistic Optimal Label Selection for Robust Prompt Learning under Partial Labels
- EVL-MCoT: Enhanced Vision-Language Multi-CoT for Harmful Meme Detection
- InterHier: Learning Interconnected Hierarchical Semantics for Open-Vocabulary Object Detection
- FedVAR: Prototype-Aligned Federated Framework for Video Anomaly Recognition
- ReHARK: Refined Hybrid Adaptive RBF Kernels for Robust One-Shot Vision-Language Adaptation
- Retrieval-augmented Prompt Learning for Pre-trained Foundation Models
- Transferable and Forecastable User Targeting Foundation Model