11 papers
EnchantDance: Unveiling the Potential of Music-Driven Dance Movement
Bo Han, Teng Zhang, Zeyu Ling +1
The task of music-driven dance generation involves creating coherent dance movements that correspond to the given music. While existing methods can produce physically plausible dan…
TransVLM: A Vision-Language Framework and Benchmark for Detecting Any Shot Transitions
Ce Chen, Yi Ren, Yuanming Li +5
Traditional Shot Boundary Detection (SBD) inherently struggles with complex transitions by formulating the task around isolated cut points, frequently yielding corrupted video shot…
Generate Your Talking Avatar from Video Reference
Zujin Guo, Zhenhui Ye, Yi Ren +4
Existing talking avatar methods typically adopt an image-to-video pipeline conditioned on a static reference image within the same scene as the target generation. This restricted,…
ReconPhys: Reconstruct Appearance and Physical Attributes from Single Video
Boyuan Wang, Xiaofeng Wang, Yongkang Li +9
Reconstructing non-rigid objects with physical plausibility remains a significant challenge. Existing approaches leverage differentiable rendering for per-scene optimization, recov…
Hierarchical Vision-Language Interaction for Facial Action Unit Detection
Yong Li, Yi Ren, Yizhe Zhang +5
Facial Action Unit (AU) detection seeks to recognize subtle facial muscle activations as defined by the Facial Action Coding System (FACS). A primary challenge w.r.t AU detection i…
InfinityHuman: Towards Long-Term Audio-Driven Human
Xiaodi Li, Pan Xie, Yi Ren +6
Audio-driven human animation has attracted wide attention thanks to its practical applications. However, critical challenges remain in generating high-resolution, long-duration vid…