activity
20232025
most citedPlug-and-Play Regulators for Image-Text Matching

45 citations · 52 across the 6 of their papers we have counts for

collaborators

6 papers

cs.CV2025

KARST: Multi-Kernel Kronecker Adaptation with Re-Scaling Transmission for Visual Classification

Yue Zhu, Haiwen Diao, Shang Gao +2

Fine-tuning pre-trained vision models for specific tasks is a common practice in computer vision. However, this process becomes more expensive as models grow larger. Recently, para…

cs.CV20245 cited

MoTrans: Customized Motion Transfer with Text-driven Video Diffusion Models

Xiaomin Li, Xu Jia, Qinghe Wang +5

Existing pretrained text-to-video (T2V) models have demonstrated impressive abilities in generating realistic videos with basic motion or camera movement. However, these models exh…

cs.AI20241 cited

LLMs Can Evolve Continually on Modality for X-Modal Reasoning

Jiazuo Yu, Haomiao Xiong, Lu Zhang +7

Multimodal Large Language Models (MLLMs) have gained significant attention due to their impressive capabilities in multimodal understanding. However, existing methods rely heavily…

cs.CV20241 cited

SHERL: Synthesizing High Accuracy and Efficient Memory for Resource-Limited Transfer Learning

Haiwen Diao, Bo Wan, Xu Jia +4

Parameter-efficient transfer learning (PETL) has emerged as a flourishing research field for adapting large pre-trained models to downstream tasks, greatly reducing trainable param…

cs.CV2024

Deep Boosting Learning: A Brand-new Cooperative Approach for Image-Text Matching

Haiwen Diao, Ying Zhang, Shang Gao +2

Image-text matching remains a challenging task due to heterogeneous semantic diversity across modalities and insufficient distance separability within triplets. Different from prev…

cs.CV202345 cited

Plug-and-Play Regulators for Image-Text Matching

Haiwen Diao, Ying Zhang, Wei Liu +2

Exploiting fine-grained correspondence and visual-semantic alignments has shown great potential in image-text matching. Generally, recent approaches first employ a cross-modal atte…