11 citations · 19 across the 4 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2021★ 4 cited
Improving Cross-Modal Alignment in Vision Language Navigation via Syntactic Information
Jialu Li, Hao Tan, Mohit Bansal
Vision language navigation is the task that requires an agent to navigate through a 3D environment based on natural language instructions. One key challenge in this task is to grou…
cs.CL2019
LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Hao Tan, Mohit Bansal
Vision-and-language reasoning requires an understanding of visual concepts, language semantics, and, most importantly, the alignment and relationships between these two modalities.…