activity
20242026
most citedDanceEditor: Towards Iterative Editable Music-driven Dance Generation with Open-Vocabulary Descriptions

1 citations · 2 across the 8 of their papers we have counts for

collaborators

9 papers

cs.CV2026

AptAvatar: Fast and Vivid Long-Form Audio-Driven Video Generation for Production-Ready Avatars

Hengyuan Zhang, Jingna Sun, Meiguang Jin +1

Production-ready audio-driven avatar generation requires efficient inference without sacrificing fidelity or motion expressiveness. However, existing acceleration methods often com…

cs.CV2026

Pantheon360: Taming Digital Twin Generation via 3D-Aware 360° Video Diffusion

Ting-Hsuan Chen, Ying-Huan Chen, Tao Tu +10

Generating complete digital twins from videos requires precise camera control, global scene coverage, and strict spatial-temporal consistency constraints that remain challenging fo…

cs.CV2026

DanceHMR: Hand-Aware Whole-Body Human Mesh Recovery from Monocular Videos

Wenhao Shen, Ming Zhou, Hengyuan Zhang +3

Monocular video human mesh recovery is essential for digital humans, avatar animation, and embodied simulation, where both temporal stability and expressive whole-body motion are r…

cs.GR20251 cited

DanceEditor: Towards Iterative Editable Music-driven Dance Generation with Open-Vocabulary Descriptions

Hengyuan Zhang, Zhe Li, Xingqun Qi +4

Generating coherent and diverse human dances from music signals has gained tremendous progress in animating virtual avatars. While existing methods support direct dance synthesis,…

cs.CV2025

CoGesture: Towards Coherent Concurrent Co-speech 3D Gesture Generation with Interactive Diffusion

Xingqun Qi, Yatian Wang, Hengyuan Zhang +6

Generating gestures from human speech has gained tremendous progress in animating virtual avatars. While the existing methods enable synthesizing gestures cooperated by individual…

cs.RO2025

RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete

Yuheng Ji, Huajie Tan, Jiayu Shi +14

Recent advancements in Multimodal Large Language Models (MLLMs) have shown remarkable capabilities across various multimodal contexts. However, their application in robotic scenari…