activity
20232025
most citedMixSpeech: Cross-Modality Self-Learning with Audio-Visual Stream Mixup for Visual Speech Translation and Recognition

8 citations · 14 across the 13 of their papers we have counts for

collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2024

Landmark-guided Diffusion Model for High-fidelity and Temporally Coherent Talking Head Generation

Jintao Tan, Xize Cheng, Lingyu Xiong +6

Audio-driven talking head generation is a significant and challenging task applicable to various fields such as virtual avatars, film production, and online conferences. However, t…

cs.CV20241 cited

OmniBind: Large-scale Omni Multimodal Representation via Binding Spaces

Zehan Wang, Ziang Zhang, Hang Zhang +5

Recently, human-computer interaction with various modalities has shown promising applications, like GPT-4o and Gemini. Given the foundational role of multimodal joint representatio…

cs.CV2023

3DRP-Net: 3D Relative Position-aware Network for 3D Visual Grounding

Zehan Wang, Haifeng Huang, Yang Zhao +5

3D visual grounding aims to localize the target object in a 3D point cloud by a free-form language description. Typically, the sentences describing the target object tend to provid…

cs.CV20231 cited

Distilling Coarse-to-Fine Semantic Matching Knowledge for Weakly Supervised 3D Visual Grounding

Zehan Wang, Haifeng Huang, Yang Zhao +5

3D visual grounding involves finding a target object in a 3D scene that corresponds to a given sentence query. Although many approaches have been proposed and achieved impressive p…

cs.CV20238 cited

MixSpeech: Cross-Modality Self-Learning with Audio-Visual Stream Mixup for Visual Speech Translation and Recognition

Xize Cheng, Linjun Li, Tao Jin +7

Multi-media communications facilitate global interaction among people. However, despite researchers exploring cross-lingual translation techniques such as machine translation and a…