787 citations · 1.8k across the 51 of their papers we have counts for
16 papers · 1 filter
Robustness Evaluation of Transformer-based Form Field Extractors via Form Attacks
Le Xue, Mingfei Gao, Zeyuan Chen +2
We propose a novel framework to evaluate the robustness of transformer-based form field extraction methods via form attacks. We introduce 14 novel form transformations to evaluate…
A Theory-Driven Self-Labeling Refinement Method for Contrastive Representation Learning
Pan Zhou, Caiming Xiong, Xiao-Tong Yuan +1
For an image query, unsupervised contrastive learning labels crops of the same image as positives, and other image crops as negatives. Although intuitive, such a native label assig…
Structured Scene Memory for Vision-Language Navigation
Hanqing Wang, Wenguan Wang, Wei Liang +2
Recently, numerous algorithms have been developed to tackle the problem of vision-language navigation (VLN), i.e., entailing an agent to navigate 3D environments through following…
MoPro: Webly Supervised Learning with Momentum Prototypes
Junnan Li, Caiming Xiong, Steven C. H. Hoi
We propose a webly-supervised representation learning method that does not suffer from the annotation unscalability of supervised learning, nor the computation unscalability of sel…
Prototypical Contrastive Learning of Unsupervised Representations
Junnan Li, Pan Zhou, Caiming Xiong +1
This paper presents Prototypical Contrastive Learning (PCL), an unsupervised representation learning method that addresses the fundamental limitations of instance-wise contrastive…
VD-BERT: A Unified Vision and Dialog Transformer with BERT
Yue Wang, Shafiq Joty, Michael R. Lyu +3
Visual dialog is a challenging vision-language task, where a dialog agent needs to answer a series of questions through reasoning on the image content and dialog history. Prior wor…