5 papers
TeMuDance: Contrastive Alignment-Based Textual Control for Music-Driven Dance Generation
Xinran Liu, Diptesh Kanojia, Wenwu Wang +1
Existing music-driven dance generation approaches have achieved strong realism and effective audio-motion alignment. However, they generally lack semantic controllability, making i…
MAGIC-Talk: Motion-aware Audio-Driven Talking Face Generation with Customizable Identity Control
Fatemeh Nazarieh, Zhenhua Feng, Diptesh Kanojia +2
Audio-driven talking face generation has gained significant attention for applications in digital media and virtual avatars. While recent methods improve audio-lip synchronization,…
DGFM: Full Body Dance Generation Driven by Music Foundation Models
Xinran Liu, Zhenhua Feng, Diptesh Kanojia +1
In music-driven dance motion generation, most existing methods use hand-crafted features and neglect that music foundation models have profoundly impacted cross-modal content gener…
PortraitTalk: Towards Customizable One-Shot Audio-to-Talking Face Generation
Fatemeh Nazarieh, Zhenhua Feng, Diptesh Kanojia +2
Audio-driven talking face generation is a challenging task in digital communication. Despite significant progress in the area, most existing methods concentrate on audio-lip synchr…
Rethinking Positive Pairs in Contrastive Learning
Jiantao Wu, Sara Atito, Zhenhua Feng +3
The training methods in AI do involve semantically distinct pairs of samples. However, their role typically is to enhance the between class separability. The actual notion of simil…