53 citations · 184 across the 18 of their papers we have counts for
28 papers
Video Autoencoder: self-supervised disentanglement of static 3D structure and motion
Zihang Lai, Sifei Liu, Alexei A. Efros +1
A video autoencoder is proposed for learning disentan- gled representations of 3D structure and camera pose from videos in a self-supervised manner. Relying on temporal continuity…
Test-Time Personalization with a Transformer for Human Pose Estimation
Yizhuo Li, Miao Hao, Zonglin Di +2
We propose to personalize a human pose estimator given a set of test images of a person without using any manual annotations. While there is a significant advancement in human pose…
Semi-Supervised 3D Hand-Object Poses Estimation with Interactions in Time
Shaowei Liu, Hanwen Jiang, Jiarui Xu +2
Estimating 3D hand and object pose from a single image is an extremely challenging problem: hands and objects are often self-occluded during interactions, and the 3D annotations ar…
Single RGB-D Camera Teleoperation for General Robotic Manipulation
Quan Vuong, Yuzhe Qin, Runlin Guo +3
We propose a teleoperation system that uses a single RGB-D camera as the human motion capture device. Our system can perform general manipulation tasks such as cloth folding, hamme…
DAIR: Disentangled Attention Intrinsic Regularization for Safe and Efficient Bimanual Manipulation
Minghao Zhang, Pingcheng Jian, Yi Wu +2
We address the problem of safely solving complex bimanual robot manipulation tasks with sparse rewards. Such challenging tasks can be decomposed into sub-tasks that are accomplisha…
Contrastive Learning of Image Representations with Cross-Video Cycle-Consistency
Haiping Wu, Xiaolong Wang
Recent works have advanced the performance of self-supervised representation learning by a large margin. The core among these methods is intra-image invariance learning. Two differ…