168 citations · 255 across the 7 of their papers we have counts for
11 papers
SeA: Semantic Adversarial Augmentation for Last Layer Features from Unsupervised Representation Learning
Qi Qian, Yuanhong Xu, Juhua Hu
Deep features extracted from certain layers of a pre-trained deep model show superior performance over the conventional hand-crafted features. Compared with fine-tuning or linear p…
Intra-Modal Proxy Learning for Zero-Shot Visual Categorization with CLIP
Qi Qian, Yuanhong Xu, Juhua Hu
Vision-language pre-training methods, e.g., CLIP, demonstrate an impressive zero-shot performance on visual categorizations with the class proxy from the text embedding of the clas…
mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Qinghao Ye, Haiyang Xu, Guohai Xu +15
Large language models (LLMs) have demonstrated impressive zero-shot abilities on a variety of open-ended tasks, while recent research has also explored the use of LLMs for multi-mo…
Improved Visual Fine-tuning with Natural Language Supervision
Junyang Wang, Yuanhong Xu, Juhua Hu +3
Fine-tuning a visual pre-trained model can leverage the semantic information from large-scale pre-training data and mitigate the over-fitting problem on downstream vision tasks wit…
mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and Video
Haiyang Xu, Qinghao Ye, Ming Yan +12
Recent years have witnessed a big convergence of language, vision, and multi-modal pretraining. In this work, we present mPLUG-2, a new unified paradigm with modularized design for…
An Empirical Study on Distribution Shift Robustness From the Perspective of Pre-Training and Data Augmentation
Ziquan Liu, Yi Xu, Yuanhong Xu +5
The performance of machine learning models under distribution shift has been the focus of the community in recent years. Most of current methods have been proposed to improve the r…