4 papers
H-Flow: Self-supervised Human Scene Flow via Physics-inspired Joint Multi-modal Learning
Zhanbo Huang, Xiaoming Liu, Yu Kong
Parametric human models capture global pose but cannot represent the non-rigid surface dynamics of clothing and soft tissue. Generic scene flow estimates dense motion but breaks do…
From 3D Pose to Prose: Biomechanics-Grounded Vision--Language Coaching
Yuyang Ji, Yixuan Shen, Shengjie Zhu +2
We present BioCoach, a biomechanics-grounded vision--language framework for fitness coaching from streaming video. BioCoach fuses visual appearance and 3D skeletal kinematics, thro…
Show Me: Unifying Instructional Image and Video Generation with Diffusion Models
Yujiang Pu, Zhanbo Huang, Vishnu Boddeti +1
Generating visual instructions in a given context is essential for developing interactive world simulators. While prior works address this problem through either text-guided image…
H-MoRe: Learning Human-centric Motion Representation for Action Analysis
Zhanbo Huang, Xiaoming Liu, Yu Kong
In this paper, we propose H-MoRe, a novel pipeline for learning precise human-centric motion representation. Our approach dynamically preserves relevant human motion while filterin…