activity
20182021
most citedL2C: Describing Visual Differences Needs Semantic Understanding of Individuals

2 citations · 2 across the 1 of their papers we have counts for

collaborators

9 papers

cs.CV20212 cited

L2C: Describing Visual Differences Needs Semantic Understanding of Individuals

An Yan, Xin Eric Wang, Tsu-Jui Fu +1

Recent advances in language and vision push forward the research of captioning a single image to describing visual differences between image pairs. Suppose there are two images, I_…

cs.CV2020

Learning to Stop: A Simple yet Effective Approach to Urban Vision-Language Navigation

Jiannan Xiang, Xin Eric Wang, William Yang Wang

Vision-and-Language Navigation (VLN) is a natural language grounding task where an agent learns to follow language instructions and navigate to specified destinations in real-world…

cs.CL2020

Towards Understanding Sample Variance in Visually Grounded Language Generation: Evaluations and Observations

Wanrong Zhu, Xin Eric Wang, Pradyumna Narayana +3

A major challenge in visually grounded language generation is to build robust benchmark datasets and models that can generalize well in real-world settings. To do this, it is criti…

cs.CL2020

Multimodal Text Style Transfer for Outdoor Vision-and-Language Navigation

Wanrong Zhu, Xin Eric Wang, Tsu-Jui Fu +5

One of the most challenging topics in Natural Language Processing (NLP) is visually-grounded language understanding and reasoning. Outdoor vision-and-language navigation (VLN) is s…

cs.AI2020

Environment-agnostic Multitask Learning for Natural Language Grounded Navigation

Xin Eric Wang, Vihan Jain, Eugene Ie +3

Recent research efforts enable study for natural language grounded navigation in photo-realistic environments, e.g., following natural language instructions or dialog. However, exi…

cs.CV2019

Unsupervised Reinforcement Learning of Transferable Meta-Skills for Embodied Navigation

Juncheng Li, Xin Wang, Siliang Tang +4

Visual navigation is a task of training an embodied agent by intelligently navigating to a target object (e.g., television) using only visual observations. A key challenge for curr…