activity
20202022
most citedAdaFuse: Adaptive Multiview Fusion for Accurate Human Pose Estimation in the Wild

113 citations · 137 across the 6 of their papers we have counts for

collaborators
Showing cs.CVShow all

10 papers · 1 filter

cs.CV20235 cited

V-DETR: DETR with Vertex Relative Position Encoding for 3D Object Detection

Yichao Shen, Zigang Geng, Yuhui Yuan +6

We introduce a highly performant 3D object detector for point clouds using the DETR framework. The prior attempts all end up with suboptimal results because they fail to learn accu…

cs.CV2022

Correlation-Aware Deep Tracking

Fei Xie, Chunyu Wang, Guangting Wang +3

Robustness and discrimination power are two fundamental requirements in visual object tracking. In most tracking paradigms, we find that the features extracted by the popular Siame…

cs.CV20216 cited

Relational Self-Attention: What's Missing in Attention for Video Understanding

Manjin Kim, Heeseung Kwon, Chunyu Wang +2

Convolution has been arguably the most important feature transform for modern neural networks, leading to the advance of deep learning. Recent emergence of Transformer networks, wh…

cs.CV20215 cited

VoxelTrack: Multi-Person 3D Human Pose Estimation and Tracking in the Wild

Yifu Zhang, Chunyu Wang, Xinggang Wang +2

We present VoxelTrack for multi-person 3D pose estimation and tracking from a few cameras which are separated by wide baselines. It employs a multi-branch network to jointly estima…

cs.CV20213 cited

Context Modeling in 3D Human Pose Estimation: A Unified Perspective

Xiaoxuan Ma, Jiajun Su, Chunyu Wang +2

Estimating 3D human pose from a single image suffers from severe ambiguity since multiple 3D joint configurations may have the same 2D projection. The state-of-the-art methods ofte…

cs.CV2020

An Empirical Study of the Collapsing Problem in Semi-Supervised 2D Human Pose Estimation

Rongchang Xie, Chunyu Wang, Wenjun Zeng +1

Semi-supervised learning aims to boost the accuracy of a model by exploring unlabeled images. The state-of-the-art methods are consistency-based which learn about unlabeled images…