6 papers
CARPRT: Class-Aware Zero-Shot Prompt Reweighting for Black-Box Vision-Language Models
Ruijiang Dong, Zesheng Ye, Jianzhong Qi +4
Pre-trained vision-language models (VLMs) enable zero-shot image classification by computing the similarity score between an image and textual descriptions, typically formed by ins…
Exploring Weak-to-Strong Generalization for CLIP-based Classification
Jinhao Li, Sarah M. Erfani, Lei Feng +2
Aligning large-scale commercial models with user intent is crucial to preventing harmful outputs. Current methods rely on human supervision but become impractical as model complexi…
Neural Network Reprogrammability: A Unified Theme on Model Reprogramming, Prompt Tuning, and Prompt Instruction
Zesheng Ye, Chengyi Cai, Ruijiang Dong +4
As large-scale pre-trained foundation models continue to expand in size and capability, efficiently adapting them to specific downstream tasks has become increasingly critical. Des…
Test-Time Multimodal Backdoor Detection by Contrastive Prompting
Yuwei Niu, Shuo He, Qi Wei +3
While multimodal contrastive learning methods (e.g., CLIP) can achieve impressive zero-shot classification performance, recent research has revealed that these methods are vulnerab…
Understanding Model Reprogramming for CLIP via Decoupling Visual Prompts
Chengyi Cai, Zesheng Ye, Lei Feng +2
Model reprogramming adapts pretrained models to downstream tasks by modifying only the input and output spaces. Visual reprogramming (VR) is one instance for vision tasks that adds…
Attribute-based Visual Reprogramming for Vision-Language Models
Chengyi Cai, Zesheng Ye, Lei Feng +2
Visual reprogramming (VR) reuses pre-trained vision models for downstream image classification tasks by adding trainable noise patterns to inputs. When applied to vision-language m…