5 papers
STAR: Skeletal Token Alignment and Rearrangement for Interaction Recognition
Yuhang Wen, Mengyuan Liu, Zixuan Tang +3
Understanding physical human-robot and human-human interactions is a challenging yet emerging topic in 3D vision. While most existing methods rely on skeleton sequences--effective…
Cross-Modal Retrieval for Motion and Text via DropTriple Loss
Sheng Yan, Yang Liu, Haoqiang Wang +3
Cross-modal retrieval of image-text and video-text is a prominent research area in computer vision and natural language processing. However, there has been insufficient attention g…
PoseMoE: Mixture-of-Experts Network for Monocular 3D Human Pose Estimation
Mengyuan Liu, Jiajie Liu, Jinyan Zhang +2
The lifting-based methods have dominated monocular 3D human pose estimation by leveraging detected 2D poses as intermediate representations. The 2D component of the final 3D human…
Heatmap Pooling Network for Action Recognition from RGB Videos
Mengyuan Liu, Jinfu Liu, Yongkang Jiang +1
Human action recognition (HAR) in videos has garnered widespread attention due to the rich information in RGB videos. Nevertheless, existing methods for extracting deep features fr…
Recognizing Actions from Robotic View for Natural Human-Robot Interaction
Ziyi Wang, Peiming Li, Hong Liu +5
Natural Human-Robot Interaction (N-HRI) requires robots to recognize human actions at varying distances and states, regardless of whether the robot itself is in motion or stationar…