45 citations · 52 across the 6 of their papers we have counts for
6 papers
KARST: Multi-Kernel Kronecker Adaptation with Re-Scaling Transmission for Visual Classification
Yue Zhu, Haiwen Diao, Shang Gao +2
Fine-tuning pre-trained vision models for specific tasks is a common practice in computer vision. However, this process becomes more expensive as models grow larger. Recently, para…
MoTrans: Customized Motion Transfer with Text-driven Video Diffusion Models
Xiaomin Li, Xu Jia, Qinghe Wang +5
Existing pretrained text-to-video (T2V) models have demonstrated impressive abilities in generating realistic videos with basic motion or camera movement. However, these models exh…
LLMs Can Evolve Continually on Modality for X-Modal Reasoning
Jiazuo Yu, Haomiao Xiong, Lu Zhang +7
Multimodal Large Language Models (MLLMs) have gained significant attention due to their impressive capabilities in multimodal understanding. However, existing methods rely heavily…
SHERL: Synthesizing High Accuracy and Efficient Memory for Resource-Limited Transfer Learning
Haiwen Diao, Bo Wan, Xu Jia +4
Parameter-efficient transfer learning (PETL) has emerged as a flourishing research field for adapting large pre-trained models to downstream tasks, greatly reducing trainable param…
Deep Boosting Learning: A Brand-new Cooperative Approach for Image-Text Matching
Haiwen Diao, Ying Zhang, Shang Gao +2
Image-text matching remains a challenging task due to heterogeneous semantic diversity across modalities and insufficient distance separability within triplets. Different from prev…
Plug-and-Play Regulators for Image-Text Matching
Haiwen Diao, Ying Zhang, Wei Liu +2
Exploiting fine-grained correspondence and visual-semantic alignments has shown great potential in image-text matching. Generally, recent approaches first employ a cross-modal atte…