activity
20242026
collaborators

5 papers

cs.RO2026

KineVLA: Towards Kinematics-Aware Vision-Language-Action Models with Bi-Level Action Decomposition

Gaoge Han, Zhengqing Gao, Ziwen Li +5

In this paper, we introduce a novel kinematics-rich vision-language-action (VLA) task, in which language commands densely encode diverse kinematic attributes (such as direction, tr…

cs.CV2026

From 2D Alignment to 3D Plausibility: Unifying Heterogeneous 2D Priors and Penetration-Free Diffusion for Occlusion-Robust Two-Hand Reconstruction

Gaoge Han, Yongkang Cheng, Zhe Chen +2

Two-hand reconstruction from monocular images is hampered by complex poses and severe occlusions, which often cause interaction misalignment and two-hand penetration. We address th…

cs.GR2025

Inter-Diffusion Generation Model of Speakers and Listeners for Effective Communication

Jinhe Huang, Yongkang Cheng, Yuming Hang +4

Full-body gestures play a pivotal role in natural interactions and are crucial for achieving effective communication. Nevertheless, most existing studies primarily focus on the ges…

cs.SD2024

Conditional GAN for Enhancing Diffusion Models in Efficient and Authentic Global Gesture Generation from Audios

Yongkang Cheng, Mingjiang Liang, Shaoli Huang +3

Audio-driven simultaneous gesture generation is vital for human-computer communication, AI games, and film production. While previous research has shown promise, there are still li…

cs.CV2024

ReinDiffuse: Crafting Physically Plausible Motions with Reinforced Diffusion Model

Gaoge Han, Mingjiang Liang, Jinglei Tang +3

Generating human motion from textual descriptions is a challenging task. Existing methods either struggle with physical credibility or are limited by the complexities of physics si…