activity
20192023
most citedUnsupervised Person Image Generation with Semantic Parsing Transformation

17 citations · 35 across the 5 of their papers we have counts for

collaborators
Showing cs.CVShow all

9 papers · 1 filter

cs.CV20234 cited

Learning and Evaluating Human Preferences for Conversational Head Generation

Mohan Zhou, Yalong Bai, Wei Zhang +3

A reliable and comprehensive evaluation metric that aligns with manual preference assessments is crucial for conversational head video synthesis methods development. Existing quant…

cs.CV2023

Deep Equilibrium Multimodal Fusion

Jinhong Ni, Yalong Bai, Wei Zhang +2

Multimodal fusion integrates the complementary information present in multiple modalities and has gained much attention recently. Most existing fusion approaches either learn a fix…

cs.CV2022

Visualizing and Understanding Patch Interactions in Vision Transformer

Jie Ma, Yalong Bai, Bineng Zhong +3

Vision Transformer (ViT) has become a leading tool in various computer vision tasks, owing to its unique self-attention mechanism that learns visual representations explicitly thro…

cs.CV20228 cited

Freeform Body Motion Generation from Speech

Jing Xu, Wei Zhang, Yalong Bai +2

People naturally conduct spontaneous body motions to enhance their speeches while giving talks. Body motion generation from speech is inherently difficult due to the non-determinis…

cs.CV20206 cited

Classes Matter: A Fine-grained Adversarial Approach to Cross-domain Semantic Segmentation

Haoran Wang, Tong Shen, Wei Zhang +2

Despite great progress in supervised semantic segmentation,a large performance drop is usually observed when deploying the model in the wild. Domain adaptation methods tackle the i…

cs.CV2020

Look-into-Object: Self-supervised Structure Modeling for Object Recognition

Mohan Zhou, Yalong Bai, Wei Zhang +2

Most object recognition approaches predominantly focus on learning discriminative visual patterns while overlooking the holistic object structure. Though important, structure model…