414 citations · 473 across the 3 of their papers we have counts for
4 papers
CLIP-Event: Connecting Text and Images with Event Structures
Manling Li, Ruochen Xu, Shuohang Wang +6
Vision-language (V+L) pretraining models have achieved great success in supporting multimedia applications by understanding the alignments between images and text. While existing v…
MLP Architectures for Vision-and-Language Modeling: An Empirical Study
Yixin Nie, Linjie Li, Zhe Gan +6
We initiate the first empirical study on the use of MLP architectures for vision-and-language (VL) fusion. Through extensive experiments on 5 VL tasks and 5 robust VQA benchmarks,…
A Compare-Aggregate Model for Matching Text Sequences
Shuohang Wang, Jing Jiang
Many NLP tasks including machine comprehension, answer selection and text entailment require the comparison between sequences. Matching the important units between sequences is a k…
Machine Comprehension Using Match-LSTM and Answer Pointer
Shuohang Wang, Jing Jiang
Machine comprehension of text is an important problem in natural language processing. A recently released dataset, the Stanford Question Answering Dataset (SQuAD), offers a large n…