most citedSpeechCraft: A Fine-grained Expressive Speech Dataset with Natural Language Description

12 citations · 38 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CV20245 cited

DanceCamAnimator: Keyframe-Based Controllable 3D Dance Camera Synthesis

Zixuan Wang, Jiayi Li, Xiaoyu Qin +4

Synthesizing camera movements from music and dance is highly challenging due to the contradicting requirements and complexities of dance cinematography. Unlike human movements, whi…

cs.SD202410 cited

VoxInstruct: Expressive Human Instruction-to-Speech Generation with Unified Multilingual Codec Language Modelling

Yixuan Zhou, Xiaoyu Qin, Zeyu Jin +5

Recent AIGC systems possess the capability to generate digital multimedia content based on human language instructions, such as text, image and video. However, when it comes to spe…

cs.MM202412 cited

SpeechCraft: A Fine-grained Expressive Speech Dataset with Natural Language Description

Zeyu Jin, Jia Jia, Qixin Wang +5

Speech-language multi-modal learning presents a significant challenge due to the fine nuanced information inherent in speech styles. Therefore, a large-scale dataset providing elab…

cs.CV20236 cited

Versatile Face Animator: Driving Arbitrary 3D Facial Avatar in RGBD Space

Haoyu Wang, Haozhe Wu, Junliang Xing +1

Creating realistic 3D facial animation is crucial for various applications in the movie production and gaming industry, especially with the burgeoning demand in the metaverse. Howe…

cs.CV20235 cited

Semantics2Hands: Transferring Hand Motion Semantics between Avatars

Zijie Ye, Jia Jia, Junliang Xing

Human hands, the primary means of non-verbal communication, convey intricate semantics in various scenarios. Due to the high sensitivity of individuals to hand motions, even minor…