60 citations · 208 across the 8 of their papers we have counts for
5 papers · 1 filter
Self-supervised Pre-training with Hard Examples Improves Visual Representations
Chunyuan Li, Xiujun Li, Lei Zhang +3
Self-supervised pre-training (SSP) employs random image transformations to generate training data for visual representation learning. In this paper, we first present a modeling fra…
VinVL: Revisiting Visual Representations in Vision-Language Models
Pengchuan Zhang, Xiujun Li, Xiaowei Hu +5
This paper presents a detailed study of improving visual representations for vision language (VL) tasks and develops an improved object detection model to provide object-centric re…
MiniVLM: A Smaller and Faster Vision-Language Model
Jianfeng Wang, Xiaowei Hu, Pengchuan Zhang +5
Recent vision-language (VL) studies have shown remarkable progress by learning generic representations from massive image-text pairs with transformer models and then fine-tuning on…
Oscar: Object-Semantics Aligned Pre-training for Vision-Language Tasks
Xiujun Li, Xi Yin, Chunyuan Li +9
Large-scale pre-training methods of learning cross-modal representations on image-text pairs are becoming popular for vision-language tasks. While existing methods simply concatena…
Towards Learning a Generic Agent for Vision-and-Language Navigation via Pre-training
Weituo Hao, Chunyuan Li, Xiujun Li +2
Learning to navigate in a visual environment following natural-language instructions is a challenging task, because the multimodal inputs to the agent are highly variable, and the…