5 papers
Collaborative Low-Rank Adaptation for Pre-Trained Vision Transformers
Zheng Liu, Jinchao Zhu, Gao Huang
Low-rank adaptation (LoRA) has achieved remarkable success in fine-tuning pre-trained vision transformers for various downstream tasks. Existing studies mainly focus on exploring m…
Progressive Confident Masking Attention Network for Audio-Visual Segmentation
Yuxuan Wang, Jinchao Zhu, Feng Dong +1
Audio and visual signals typically occur simultaneously, and humans possess an innate ability to correlate and synchronize information from these two modalities. Recently, a challe…
Multiple-Exit Tuning: Towards Inference-Efficient Adaptation for Vision Transformer
Zheng Liu, Jinchao Zhu, Nannan Li +1
Parameter-efficient transfer learning (PETL) has shown great potential in adapting a vision transformer (ViT) pre-trained on large-scale datasets to various downstream tasks. Exist…
A-SDM: Accelerating Stable Diffusion through Model Assembly and Feature Inheritance Strategies
Jinchao Zhu, Yuxuan Wang, Siyuan Pan +3
The Stable Diffusion Model (SDM) is a prevalent and effective model for text-to-image (T2I) and image-to-image (I2I) generation. Despite various attempts at sampler optimization, m…
A-SDM: Accelerating Stable Diffusion through Redundancy Removal and Performance Optimization
Jinchao Zhu, Yuxuan Wang, Xiaobing Tu +3
The Stable Diffusion Model (SDM) is a popular and efficient text-to-image (t2i) generation and image-to-image (i2i) generation model. Although there have been some attempts to redu…