3 papers
cs.CV2026
From 3D Pose to Prose: Biomechanics-Grounded Vision--Language Coaching
Yuyang Ji, Yixuan Shen, Shengjie Zhu +2
We present BioCoach, a biomechanics-grounded vision--language framework for fitness coaching from streaming video. BioCoach fuses visual appearance and 3D skeletal kinematics, thro…
cs.CV2025
Show Me: Unifying Instructional Image and Video Generation with Diffusion Models
Yujiang Pu, Zhanbo Huang, Vishnu Boddeti +1
Generating visual instructions in a given context is essential for developing interactive world simulators. While prior works address this problem through either text-guided image…
cs.CV2025
H-MoRe: Learning Human-centric Motion Representation for Action Analysis
Zhanbo Huang, Xiaoming Liu, Yu Kong
In this paper, we propose H-MoRe, a novel pipeline for learning precise human-centric motion representation. Our approach dynamically preserves relevant human motion while filterin…