214 citations · 633 across the 33 of their papers we have counts for
9 papers · 1 filter
Attentive Mask CLIP
Yifan Yang, Weiquan Huang, Yixuan Wei +8
Image token removal is an efficient augmentation strategy for reducing the cost of computing image features. However, this efficient augmentation strategy has been found to adverse…
Two-Stream Network for Sign Language Recognition and Translation
Yutong Chen, Ronglai Zuo, Fangyun Wei +3
Sign languages are visual languages using manual articulations and non-manual elements to convey information. For sign language recognition and translation, the majority of existin…
AniFaceGAN: Animatable 3D-Aware Face Image Generation for Video Avatars
Yue Wu, Yu Deng, Jiaolong Yang +3
Although 2D generative models have made great progress in face image generation and animation, they often suffer from undesirable artifacts such as 3D inconsistency when rendering…
Conditional DETR V2: Efficient Detection Transformer with Box Queries
Xiaokang Chen, Fangyun Wei, Gang Zeng +1
In this paper, we are interested in Detection Transformer (DETR), an end-to-end object detection approach based on a transformer encoder-decoder architecture without hand-crafted p…
Boosting Zero-shot Learning via Contrastive Optimization of Attribute Representations
Yu Du, Miaojing Shi, Fangyun Wei +1
Zero-shot learning (ZSL) aims to recognize classes that do not have samples in the training set. One representative solution is to directly learn an embedding function associating…
Unsupervised Prompt Learning for Vision-Language Models
Tony Huang, Jack Chu, Fangyun Wei
Contrastive vision-language models like CLIP have shown great progress in transfer learning. In the inference stage, the proper text description, also known as prompt, needs to be…