7 citations · 7 across the 5 of their papers we have counts for
7 papers
MaskSem: Semantic-Guided Masking for Learning 3D Hybrid High-Order Motion Representation
Wei Wei, Shaojie Zhang, Yonghao Dang +1
Human action recognition is a crucial task for intelligent robotics, particularly within the context of human-robot collaboration research. In self-supervised skeleton-based action…
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs
Shaojie Zhang, Jiahui Yang, Jianqin Yin +2
Multimodal Large Language Models (MLLMs) have demonstrated significant success in visual understanding tasks. However, challenges persist in adapting these models for video compreh…
A Generically Contrastive Spatiotemporal Representation Enhancement for 3D Skeleton Action Recognition
Shaojie Zhang, Jianqin Yin, Yonghao Dang
Skeleton-based action recognition is a central task in computer vision and human-robot interaction. However, most previous methods suffer from overlooking the explicit exploitation…
SiT-MLP: A Simple MLP with Point-wise Topology Feature Learning for Skeleton-based Action Recognition
Shaojie Zhang, Jianqin Yin, Yonghao Dang +1
Graph convolution networks (GCNs) have achieved remarkable performance in skeleton-based action recognition. However, previous GCN-based methods rely on elaborate human priors exce…
An Improved Baseline Framework for Pose Estimation Challenge at ECCV 2022 Visual Perception for Navigation in Human Environments Workshop
Jiajun Fu, Yonghao Dang, Ruoqi Yin +4
This technical report describes our first-place solution to the pose estimation challenge at ECCV 2022 Visual Perception for Navigation in Human Environments Workshop. In this chal…
Kinematics Modeling Network for Video-based Human Pose Estimation
Yonghao Dang, Jianqin Yin, Shaojie Zhang +2
Estimating human poses from videos is critical in human-computer interaction. Joints cooperate rather than move independently during human movement. There are both spatial and temp…