9 papers
LiveGesture Streamable Co-Speech Gesture Generation Model
Muhammad Usama Saleem, Mayur Jagdishbhai Patel, Ekkasit Pinyoanuntapong +6
We propose LiveGesture, the first fully streamable, speech-driven full-body gesture generation framework that operates with zero look-ahead and supports arbitrary sequence length.…
MAGE: Modality-Agnostic Music Generation and Target-Source Extraction
Muhammad Usama Saleem, Tejasvi Ravi, Tianyu Xu +4
Recent advances in multimodal audio generation have enabled music synthesis from text, visual cues, and other high-level conditions. However, most systems are designed for a single…
Monocular Models are Strong Learners for Multi-View Human Mesh Recovery
Haoyu Xie, Shengkai Xu, Cheng Guo +6
Multi-view human mesh recovery (HMR) is broadly deployed in diverse domains where high accuracy and strong generalization are essential. Existing approaches can be broadly grouped…
Walk Before You Dance: High-fidelity and Editable Dance Synthesis via Generative Masked Motion Prior
Foram N Shah, Parshwa Shah, Muhammad Usama Saleem +4
Recent advances in dance generation have enabled the automatic synthesis of 3D dance motions. However, existing methods still face significant challenges in simultaneously achievin…
BioPose: Biomechanically-accurate 3D Pose Estimation from Monocular Videos
Farnoosh Koleini, Muhammad Usama Saleem, Pu Wang +3
Recent advancements in 3D human pose estimation from single-camera images and videos have relied on parametric models, like SMPL. However, these models oversimplify anatomical stru…
GenHMR: Generative Human Mesh Recovery
Muhammad Usama Saleem, Ekkasit Pinyoanuntapong, Pu Wang +3
Human mesh recovery (HMR) is crucial in many computer vision applications; from health to arts and entertainment. HMR from monocular images has predominantly been addressed by dete…