8 papers · 1 filter
Generate Your Talking Avatar from Video Reference
Zujin Guo, Zhenhui Ye, Yi Ren +4
Existing talking avatar methods typically adopt an image-to-video pipeline conditioned on a static reference image within the same scene as the target generation. This restricted,…
ReconPhys: Reconstruct Appearance and Physical Attributes from Single Video
Boyuan Wang, Xiaofeng Wang, Yongkang Li +9
Reconstructing non-rigid objects with physical plausibility remains a significant challenge. Existing approaches leverage differentiable rendering for per-scene optimization, recov…
Hierarchical Vision-Language Interaction for Facial Action Unit Detection
Yong Li, Yi Ren, Yizhe Zhang +5
Facial Action Unit (AU) detection seeks to recognize subtle facial muscle activations as defined by the Facial Action Coding System (FACS). A primary challenge w.r.t AU detection i…
InfinityHuman: Towards Long-Term Audio-Driven Human
Xiaodi Li, Pan Xie, Yi Ren +6
Audio-driven human animation has attracted wide attention thanks to its practical applications. However, critical challenges remain in generating high-resolution, long-duration vid…
Beyond Overfitting: Doubly Adaptive Dropout for Generalizable AU Detection
Yong Li, Yi Ren, Xuesong Niu +3
Facial Action Units (AUs) are essential for conveying psychological states and emotional expressions. While automatic AU detection systems leveraging deep learning have progressed,…
HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation
Qijun Gan, Yi Ren, Chen Zhang +6
Human motion video generation has advanced significantly, while existing methods still struggle with accurately rendering detailed body parts like hands and faces, especially in lo…