2 papers
cs.CV2026
GA2-CLIP: Generic Attribute Anchor for Efficient Prompt Tuningin Video-Language Models
Bin Wang, Ruotong Hu, Wentong Li +5
Visual and textual soft prompt tuning can effectively improve the adaptability of Vision-Language Models (VLMs) in downstream tasks. However, fine-tuning on video tasks impairs the…
cs.CV2025
TDS-CLIP: Temporal Difference Side Network for Efficient VideoAction Recognition
Bin Wang, Wentong Li, Wenqian Wang +3
Recently, large-scale pre-trained vision-language models (e.g., CLIP), have garnered significant attention thanks to their powerful representative capabilities. This inspires resea…