activity
20152021
most citedTowards Good Practices for Very Deep Two-Stream ConvNets

385 citations · 449 across the 6 of their papers we have counts for

collaborators

9 papers

cs.CV202140 cited

Learning Dynamical Human-Joint Affinity for 3D Pose Estimation in Videos

Junhao Zhang, Yali Wang, Zhipeng Zhou +3

Graph Convolution Network (GCN) has been successfully used for 3D human pose estimation in videos. However, it is often built on the fixed human-joint affinity, according to human…

cs.CV2021

SSCAP: Self-supervised Co-occurrence Action Parsing for Unsupervised Temporal Action Segmentation

Zhe Wang, Hao Chen, Xinyu Li +4

Temporal action segmentation is a task to classify each frame in the video with an action label. However, it is quite expensive to annotate every frame in a large corpus of videos…

cs.CV20211 cited

PC-HMR: Pose Calibration for 3D Human Mesh Recovery from 2D Images/Videos

Tianyu Luan, Yali Wang, Junhao Zhang +3

The end-to-end Human Mesh Recovery (HMR) approach has been successfully used for 3D body reconstruction. However, most HMR-based frameworks reconstruct human body by directly learn…

cs.CV20206 cited

Predicting Camera Viewpoint Improves Cross-dataset Generalization for 3D Human Pose Estimation

Zhe Wang, Daeyun Shin, Charless C. Fowlkes

Monocular estimation of 3d human pose has attracted increased attention with the availability of large ground-truth motion capture datasets. However, the diversity of training data…

cs.CV201810 cited

Structured Triplet Learning with POS-tag Guided Attention for Visual Question Answering

Zhe Wang, Xiaoyi Liu, Liangjian Chen +4

Visual question answering (VQA) is of significant interest due to its potential to be a strong test of image understanding systems and to probe the connection between language and…

cs.CV2017

Temporal Segment Networks for Action Recognition in Videos

Limin Wang, Yuanjun Xiong, Zhe Wang +4

Deep convolutional networks have achieved great success for image recognition. However, for action recognition in videos, their advantage over traditional methods is not so evident…