activity
20242026
most citedConditional GAN for Enhancing Diffusion Models in Efficient and Authentic Global Gesture Generation from Audios

1 citations · 1 across the 3 of their papers we have counts for

collaborators

6 papers

cs.RO2026

KineVLA: Towards Kinematics-Aware Vision-Language-Action Models with Bi-Level Action Decomposition

Gaoge Han, Zhengqing Gao, Ziwen Li +5

In this paper, we introduce a novel kinematics-rich vision-language-action (VLA) task, in which language commands densely encode diverse kinematic attributes (such as direction, tr…

cs.GR2025

Inter-Diffusion Generation Model of Speakers and Listeners for Effective Communication

Jinhe Huang, Yongkang Cheng, Yuming Hang +4

Full-body gestures play a pivotal role in natural interactions and are crucial for achieving effective communication. Nevertheless, most existing studies primarily focus on the ges…

cs.CV2025

From 2D Alignment to 3D Plausibility: Unifying Heterogeneous 2D Priors and Penetration-Free Diffusion for Occlusion-Robust Two-Hand Reconstruction

Gaoge Han, Yongkang Cheng, Zhe Chen +2

Two-hand reconstruction from monocular images is hampered by complex poses and severe occlusions, which often cause interaction misalignment and two-hand penetration. We address th…

cs.SD20241 cited

Conditional GAN for Enhancing Diffusion Models in Efficient and Authentic Global Gesture Generation from Audios

Yongkang Cheng, Mingjiang Liang, Shaoli Huang +3

Audio-driven simultaneous gesture generation is vital for human-computer communication, AI games, and film production. While previous research has shown promise, there are still li…

cs.CV2024

ReinDiffuse: Crafting Physically Plausible Motions with Reinforced Diffusion Model

Gaoge Han, Mingjiang Liang, Jinglei Tang +3

Generating human motion from textual descriptions is a challenging task. Existing methods either struggle with physical credibility or are limited by the complexities of physics si…

cs.SD2024

ExpGest: Expressive Speaker Generation Using Diffusion Model and Hybrid Audio-Text Guidance

Yongkang Cheng, Mingjiang Liang, Shaoli Huang +3

Existing gesture generation methods primarily focus on upper body gestures based on audio features, neglecting speech content, emotion, and locomotion. These limitations result in…