5 papers
FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision
Tongyan Wang, Zhengyuan Li, Muhan Lin +5
Text-conditioned human motion generation has made rapid progress with the emergence of large-scale motion--language datasets. However, even datasets with rich long-form description…
PoseShield: Neural Collision Fields for Human Self-Collision Resolution
Zhengyuan Li, Zeyun Deng, Yifan Shen +7
Self-collision remains a persistent challenge in SMPL-based human pose estimation and motion generation. Under extreme articulations or stochastic motion synthesis, generated meshe…
MDD: A Dataset for Text-and-Music Conditioned Duet Dance Generation
Prerit Gupta, Jason Alexander Fotso-Puepi, Zhengyuan Li +2
We introduce Multimodal DuetDance (MDD), a diverse multimodal benchmark dataset designed for text-controlled and music-conditioned 3D duet dance motion generation. Our dataset comp…
SimMotionEdit: Text-Based Human Motion Editing with Motion Similarity Prediction
Zhengyuan Li, Kai Cheng, Anindita Ghosh +3
Text-based 3D human motion editing is a critical yet challenging task in computer vision and graphics. While training-free approaches have been explored, the recent release of the…
EfficientEQA: An Efficient Approach to Open-Vocabulary Embodied Question Answering
Kai Cheng, Zhengyuan Li, Xingpeng Sun +3
Embodied Question Answering (EQA) is an essential yet challenging task for robot assistants. Large vision-language models (VLMs) have shown promise for EQA, but existing approaches…