7 papers
InfinityHuman: Towards Long-Term Audio-Driven Human
Xiaodi Li, Pan Xie, Yi Ren +6
Audio-driven human animation has attracted wide attention thanks to its practical applications. However, critical challenges remain in generating high-resolution, long-duration vid…
UniTalker: Conversational Speech-Visual Synthesis
Yifan Hu, Rui Liu, Yi Ren +2
Conversational Speech Synthesis (CSS) is a key task in the user-agent interaction area, aiming to generate more expressive and empathetic speech for users. However, it is well-know…
Chain-Talker: Chain Understanding and Rendering for Empathetic Conversational Speech Synthesis
Yifan Hu, Rui Liu, Yi Ren +2
Conversational Speech Synthesis (CSS) aims to align synthesized speech with the emotional and stylistic context of user-agent interactions to achieve empathy. Current generative CS…
MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis
Ziyue Jiang, Yi Ren, Ruiqi Li +11
While recent zero-shot text-to-speech (TTS) models have significantly improved speech quality and expressiveness, mainstream systems still suffer from issues related to speech-text…
Beyond Overfitting: Doubly Adaptive Dropout for Generalizable AU Detection
Yong Li, Yi Ren, Xuesong Niu +3
Facial Action Units (AUs) are essential for conveying psychological states and emotional expressions. While automatic AU detection systems leveraging deep learning have progressed,…
HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation
Qijun Gan, Yi Ren, Chen Zhang +6
Human motion video generation has advanced significantly, while existing methods still struggle with accurately rendering detailed body parts like hands and faces, especially in lo…