activity
20182021
most citedHERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training

51 citations · 287 across the 18 of their papers we have counts for

collaborators

49 papers

cs.CV202138 cited

VALUE: A Multi-Task Benchmark for Video-and-Language Understanding Evaluation

Linjie Li, Jie Lei, Zhe Gan +12

Most existing video-and-language (VidL) research focuses on a single dataset, or multiple datasets of a single task. In reality, a truly useful VidL system is expected to be easily…

cs.CV20212 cited

Adversarial VQA: A New Benchmark for Evaluating the Robustness of VQA Models

Linjie Li, Jie Lei, Zhe Gan +1

Benefiting from large-scale pre-training, we have witnessed significant performance boost on the popular Visual Question Answering (VQA) task. Despite rapid progress, it remains un…

cs.CV20217 cited

CUPID: Adaptive Curation of Pre-training Data for Video-and-Language Representation Learning

Luowei Zhou, Jingjing Liu, Yu Cheng +2

This work concerns video-language pre-training and representation learning. In this now ubiquitous training scheme, a model first performs pre-training on paired videos and text (e…

cs.CL20216 cited

LightningDOT: Pre-training Visual-Semantic Embeddings for Real-Time Image-Text Retrieval

Siqi Sun, Yen-Chun Chen, Linjie Li +3

Multimodal pre-training has propelled great advancement in vision-and-language research. These large-scale pre-trained models, although successful, fatefully suffer from slow infer…

cs.CV20217 cited

UC2: Universal Cross-lingual Cross-modal Vision-and-Language Pre-training

Mingyang Zhou, Luowei Zhou, Shuohang Wang +4

Vision-and-language pre-training has achieved impressive success in learning multimodal representations between vision and language. To generalize this success to non-English langu…

cs.LG20219 cited

Adversarial Feature Augmentation and Normalization for Visual Recognition

Tianlong Chen, Yu Cheng, Zhe Gan +4

Recent advances in computer vision take advantage of adversarial data augmentation to ameliorate the generalization ability of classification models. Here, we present an effective…