3 papers
cs.CV2026
On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning
Zihan Zhang, Jie Hong, Siyuan Fan +2
Audio-visual Generalized Zero-shot Learning (AV-GZSL) is a challenging task that aims to classify both seen and unseen objects or scenes by integrating data from audio and visual m…
cs.CV2025
3D Human Interaction Generation: A Survey
Siyuan Fan, Wenke Huang, Xiantao Cai +1
3D human interaction generation has emerged as a key research area, focusing on producing dynamic and contextually relevant interactions between humans and various interactive enti…
cs.CV2024
TextIM: Part-aware Interactive Motion Synthesis from Text
Siyuan Fan, Bo Du, Xiantao Cai +2
In this work, we propose TextIM, a novel framework for synthesizing TEXT-driven human Interactive Motions, with a focus on the precise alignment of part-level semantics. Existing m…