25 citations · 33 across the 2 of their papers we have counts for
3 papers
cs.CV2021★ 25 cited
Seeing Out of tHe bOx: End-to-End Pre-training for Vision-Language Representation Learning
Zhicheng Huang, Zhaoyang Zeng, Yupan Huang +3
We study joint learning of Convolutional Neural Network (CNN) and Transformer for vision-language pre-training (VLPT) which aims to learn cross-modal alignments from millions of im…
cs.CV2020
Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers
Zhicheng Huang, Zhaoyang Zeng, Bei Liu +2
We propose Pixel-BERT to align image pixels with text by deep multi-modal transformers that jointly learn visual and language embedding in a unified end-to-end framework. We aim to…
cs.CV2019★ 8 cited
Learning Rich Image Region Representation for Visual Question Answering
Bei Liu, Zhicheng Huang, Zhaoyang Zeng +2
We propose to boost VQA by leveraging more powerful feature extractors by improving the representation ability of both visual and text features and the ensemble of models. For visu…