collaborators

7 papers

cs.CV2025

InfinityHuman: Towards Long-Term Audio-Driven Human

Xiaodi Li, Pan Xie, Yi Ren +6

Audio-driven human animation has attracted wide attention thanks to its practical applications. However, critical challenges remain in generating high-resolution, long-duration vid…

eess.AS2025

UniTalker: Conversational Speech-Visual Synthesis

Yifan Hu, Rui Liu, Yi Ren +2

Conversational Speech Synthesis (CSS) is a key task in the user-agent interaction area, aiming to generate more expressive and empathetic speech for users. However, it is well-know…

cs.SD2025

Chain-Talker: Chain Understanding and Rendering for Empathetic Conversational Speech Synthesis

Yifan Hu, Rui Liu, Yi Ren +2

Conversational Speech Synthesis (CSS) aims to align synthesized speech with the emotional and stylistic context of user-agent interactions to achieve empathy. Current generative CS…

eess.AS2025

MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Ziyue Jiang, Yi Ren, Ruiqi Li +11

While recent zero-shot text-to-speech (TTS) models have significantly improved speech quality and expressiveness, mainstream systems still suffer from issues related to speech-text…

cs.CV2025

Beyond Overfitting: Doubly Adaptive Dropout for Generalizable AU Detection

Yong Li, Yi Ren, Xuesong Niu +3

Facial Action Units (AUs) are essential for conveying psychological states and emotional expressions. While automatic AU detection systems leveraging deep learning have progressed,…

cs.CV2025

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation

Qijun Gan, Yi Ren, Chen Zhang +6

Human motion video generation has advanced significantly, while existing methods still struggle with accurately rendering detailed body parts like hands and faces, especially in lo…