most citedRetouchdown: Adding Touchdown to StreetLearn as a Shareable Resource for Language Grounding Tasks in Street View

20 citations · 36 across the 3 of their papers we have counts for

collaborators

5 papers

cs.LG2020

Momentum Improves Normalized SGD

Ashok Cutkosky, Harsh Mehta

We provide an improved analysis of normalized SGD showing that adding momentum provably removes the need for large batch sizes on non-convex objectives. Then, we consider the case…

cs.CV202020 cited

Retouchdown: Adding Touchdown to StreetLearn as a Shareable Resource for Language Grounding Tasks in Street View

Harsh Mehta, Yoav Artzi, Jason Baldridge +2

The Touchdown dataset (Chen et al., 2019) provides instructions by human annotators for navigation through New York City streets and for resolving spatial descriptions at a given l…

cs.LG20197 cited

VALAN: Vision and Language Agent Navigation

Larry Lansing, Vihan Jain, Harsh Mehta +2

VALAN is a lightweight and scalable software framework for deep reinforcement learning based on the SEED RL architecture. The framework facilitates the development and evaluation o…

cs.CV2019

Transferable Representation Learning in Vision-and-Language Navigation

Haoshuo Huang, Vihan Jain, Harsh Mehta +4

Vision-and-Language Navigation (VLN) tasks such as Room-to-Room (R2R) require machine agents to interpret natural language instructions and learn to act in visually realistic envir…

cs.CL20199 cited

Multi-modal Discriminative Model for Vision-and-Language Navigation

Haoshuo Huang, Vihan Jain, Harsh Mehta +2

Vision-and-Language Navigation (VLN) is a natural language grounding task where agents have to interpret natural language instructions in the context of visual scenes in a dynamic…