28 citations · 28 across the 1 of their papers we have counts for
3 papers
cs.CV2019
UNITER: UNiversal Image-TExt Representation Learning
Yen-Chun Chen, Linjie Li, Licheng Yu +5
Joint image-text embedding is the bedrock for most Vision-and-Language (V+L) tasks, where multimodality inputs are simultaneously processed for joint visual and textual understandi…
cs.CV2019★ 28 cited
Multi-step Reasoning via Recurrent Dual Attention for Visual Dialog
Zhe Gan, Yu Cheng, Ahmed El Kholy +3
This paper presents a new model for visual dialog, Recurrent Dual Attention Network (ReDAN), using multi-step reasoning to answer a series of questions about an image. In each ques…
cs.CL2016
Egyptian Arabic to English Statistical Machine Translation System for NIST OpenMT'2015
Hassan Sajjad, Nadir Durrani, Francisco Guzman +6
The paper describes the Egyptian Arabic-to-English statistical machine translation (SMT) system that the QCRI-Columbia-NYUAD (QCN) group submitted to the NIST OpenMT'2015 competiti…