11 citations · 14 across the 7 of their papers we have counts for
5 papers · 1 filter
A Bridge from Audio to Video: Phoneme-Viseme Alignment Allows Every Face to Speak Multiple Languages
Zibo Su, Kun Wei, Jiahua Li +3
Speech-driven talking face synthesis (TFS) focuses on generating lifelike facial animations from speech input. Current TFS models perform well in English but struggle with non-Engl…
Perceive, Verify and Understand Long Video: Multi-Granular Perception and Active Verification via Interactive Agents
Jiahua Li, Zhanhe Zhang, Chenghao Xu +4
Long videos, characterized by temporal complexity and sparse task-relevant information, pose significant reasoning challenges for AI systems. Although existing Large Language Model…
Do You Guys Want to Dance: Zero-Shot Compositional Human Dance Generation with Multiple Persons
Zhe Xu, Kun Wei, Xu Yang +1
Human dance generation (HDG) aims to synthesize realistic videos from images and sequences of driving poses. Despite great success, existing methods are limited to generating video…
Doubly Contrastive Deep Clustering
Zhiyuan Dang, Cheng Deng, Xu Yang +1
Deep clustering successfully provides more effective features than conventional ones and thus becomes an important technique in current unsupervised learning. However, most deep cl…
Incremental Embedding Learning via Zero-Shot Translation
Kun Wei, Cheng Deng, Xu Yang +1
Modern deep learning methods have achieved great success in machine learning and computer vision fields by learning a set of pre-defined datasets. Howerver, these methods perform u…