5 papers · 1 filter
AnchorOPT: Towards Optimizing Dynamic Anchors for Adaptive Prompt Learning
Zheng Li, Yibing Song, Xin Zhang +3
Existing prompt learning methods, which are built upon CLIP models, leverage textual tokens as anchors to guide the learnable soft tokens. This guidance improves CLIP generalizatio…
Representation Entanglement for Generation: Training Diffusion Transformers Is Much Easier Than You Think
Ge Wu, Shen Zhang, Ruijing Shi +9
REPA and its variants effectively mitigate training challenges in diffusion models by incorporating external visual representations from pretrained models, through alignment betwee…
Advancing Textual Prompt Learning with Anchored Attributes
Zheng Li, Yibing Song, Ming-Ming Cheng +2
Textual-based prompt learning methods primarily employ multiple learnable soft prompts and hard class tokens in a cascading manner as text inputs, aiming to align image and text (c…
Cascade Prompt Learning for Vision-Language Model Adaptation
Ge Wu, Xin Zhang, Zheng Li +4
Prompt learning has surfaced as an effective approach to enhance the performance of Vision-Language Models (VLMs) like CLIP when applied to downstream tasks. However, current learn…
Revisiting Prompt Pretraining of Vision-Language Models
Zhenyuan Chen, Lingfeng Yang, Shuo Chen +3
Prompt learning is an effective method to customize Vision-Language Models (VLMs) for various downstream tasks, involving tuning very few parameters of input prompt tokens. Recentl…