25 citations · 27 across the 2 of their papers we have counts for
3 papers
cs.CV2022★ 2 cited
Weakly-supervised Pre-training for 3D Human Pose Estimation via Perspective Knowledge
Zhongwei Qiu, Kai Qiu, Jianlong Fu +1
Modern deep learning-based 3D pose estimation approaches require plenty of 3D pose annotations. However, existing 3D datasets lack diversity, which limits the performance of curren…
cs.CV2021★ 25 cited
Seeing Out of tHe bOx: End-to-End Pre-training for Vision-Language Representation Learning
Zhicheng Huang, Zhaoyang Zeng, Yupan Huang +3
We study joint learning of Convolutional Neural Network (CNN) and Transformer for vision-language pre-training (VLPT) which aims to learn cross-modal alignments from millions of im…
cs.CV2020
Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers
Zhicheng Huang, Zhaoyang Zeng, Bei Liu +2
We propose Pixel-BERT to align image pixels with text by deep multi-modal transformers that jointly learn visual and language embedding in a unified end-to-end framework. We aim to…