17 citations · 31 across the 4 of their papers we have counts for
7 papers
Visualizing and Understanding Patch Interactions in Vision Transformer
Jie Ma, Yalong Bai, Bineng Zhong +3
Vision Transformer (ViT) has become a leading tool in various computer vision tasks, owing to its unique self-attention mechanism that learns visual representations explicitly thro…
Freeform Body Motion Generation from Speech
Jing Xu, Wei Zhang, Yalong Bai +2
People naturally conduct spontaneous body motions to enhance their speeches while giving talks. Body motion generation from speech is inherently difficult due to the non-determinis…
Classes Matter: A Fine-grained Adversarial Approach to Cross-domain Semantic Segmentation
Haoran Wang, Tong Shen, Wei Zhang +2
Despite great progress in supervised semantic segmentation,a large performance drop is usually observed when deploying the model in the wild. Domain adaptation methods tackle the i…
Look-into-Object: Self-supervised Structure Modeling for Object Recognition
Mohan Zhou, Yalong Bai, Wei Zhang +2
Most object recognition approaches predominantly focus on learning discriminative visual patterns while overlooking the holistic object structure. Though important, structure model…
Down to the Last Detail: Virtual Try-on with Detail Carving
Jiahang Wang, Wei Zhang, Weizhong Liu +1
Virtual try-on under arbitrary poses has attracted lots of research attention due to its huge potential applications. However, existing methods can hardly preserve the details in c…
Everyone is a Cartoonist: Selfie Cartoonization with Attentive Adversarial Networks
Xinyu Li, Wei Zhang, Tong Shen +1
Selfie and cartoon are two popular artistic forms that are widely presented in our daily life. Despite the great progress in image translation/stylization, few techniques focus spe…