17 citations · 35 across the 5 of their papers we have counts for
9 papers · 1 filter
Learning and Evaluating Human Preferences for Conversational Head Generation
Mohan Zhou, Yalong Bai, Wei Zhang +3
A reliable and comprehensive evaluation metric that aligns with manual preference assessments is crucial for conversational head video synthesis methods development. Existing quant…
Deep Equilibrium Multimodal Fusion
Jinhong Ni, Yalong Bai, Wei Zhang +2
Multimodal fusion integrates the complementary information present in multiple modalities and has gained much attention recently. Most existing fusion approaches either learn a fix…
Visualizing and Understanding Patch Interactions in Vision Transformer
Jie Ma, Yalong Bai, Bineng Zhong +3
Vision Transformer (ViT) has become a leading tool in various computer vision tasks, owing to its unique self-attention mechanism that learns visual representations explicitly thro…
Freeform Body Motion Generation from Speech
Jing Xu, Wei Zhang, Yalong Bai +2
People naturally conduct spontaneous body motions to enhance their speeches while giving talks. Body motion generation from speech is inherently difficult due to the non-determinis…
Classes Matter: A Fine-grained Adversarial Approach to Cross-domain Semantic Segmentation
Haoran Wang, Tong Shen, Wei Zhang +2
Despite great progress in supervised semantic segmentation,a large performance drop is usually observed when deploying the model in the wild. Domain adaptation methods tackle the i…
Look-into-Object: Self-supervised Structure Modeling for Object Recognition
Mohan Zhou, Yalong Bai, Wei Zhang +2
Most object recognition approaches predominantly focus on learning discriminative visual patterns while overlooking the holistic object structure. Though important, structure model…