42 citations · 44 across the 5 of their papers we have counts for
5 papers
Efficient and Long-Tailed Generalization for Pre-trained Vision-Language Model
Jiang-Xin Shi, Chi Zhang, Tong Wei +1
Pre-trained vision-language models like CLIP have shown powerful zero-shot inference ability via image-text matching and prove to be strong few-shot learners in various downstream…
Offline Imitation Learning with Model-based Reverse Augmentation
Jie-Jing Shao, Hao-Sen Shi, Lan-Zhe Guo +1
In offline Imitation Learning (IL), one of the main challenges is the \textit{covariate shift} between the expert observations and the actual distribution encountered by the agent,…
Investigating the Limitation of CLIP Models: The Worst-Performing Categories
Jie-Jing Shao, Jiang-Xin Shi, Xiao-Wen Yang +2
Contrastive Language-Image Pre-training (CLIP) provides a foundation model by integrating natural language into visual concepts, enabling zero-shot recognition on downstream tasks.…
Learning Image Deraining Transformer Network with Dynamic Dual Self-Attention
Zhentao Fan, Hongming Chen, Yufeng Li
Recently, Transformer-based architecture has been introduced into single image deraining task due to its advantage in modeling non-local information. However, existing approaches t…
USB: A Unified Semi-supervised Learning Benchmark for Classification
Yidong Wang, Hao Chen, Yue Fan +19
Semi-supervised learning (SSL) improves model generalization by leveraging massive unlabeled data to augment limited labeled samples. However, currently, popular SSL evaluation pro…