1 paper
Bin Wang, Ruotong Hu, Wentong Li +5
Visual and textual soft prompt tuning can effectively improve the adaptability of Vision-Language Models (VLMs) in downstream tasks. However, fine-tuning on video tasks impairs the…