18 papers
VideoAgent: All-in-One Framework for Video Understanding and Editing
Hengji Zhou, Lingxuan Huang, Jian Wang +4
Video editing has become essential in digital media creation, yet existing automated systems are restricted to short segment processing and domain-specific tasks. They face two cri…
Motion-Focused Latent Action Enables Cross-Embodiment VLA Training from Human EgoVideos
Runze Xu, Yiluo Zhang, Jian Wang +2
Training generalist Vision-Language-Action(VLA) models typically requires massive, diverse robotic datasets with high-fidelity action annotations. While egocentric human manipulati…
RAGA: Real Time Ray Traced Gaussian Shadow Casting for 3DGS Avatar-Scene Interaction
Aymen Mir, Riza Alp Guler, Jian Wang +3
We study the problem of physically plausible shadow casting when animating 3D Gaussian Splatting (3DGS) avatars, either individually or in multi-avatar and object-interaction scena…
CapTalk: Text-Guided Stylization and Speech-Driven 3D Head Animation
Xuangeng Chu, Yuan Gan, Ziteng Cui +4
Audio-driven 3D facial animation aims to generate synchronized lip movements and vivid facial expressions from arbitrary audio clips. While existing methods can produce synchronize…
Helix4D: Complex 4D Mesh Generation
Jiraphon Yenphraphai, Jianqi Chen, Jian Wang +6
Current video-to-4D methods struggle with complex topology changes, transparent materials, thin structures, and inner surfaces. We present Helix4D, a dynamic mesh generation framew…
HandX: Scaling Bimanual Motion and Interaction Generation
Zimu Zhang, Yucheng Zhang, Xiyan Xu +8
Synthesizing human motion has advanced rapidly, yet realistic hand motion and bimanual interaction remain underexplored. Whole-body models often miss the fine-grained cues that dri…