215 citations · 259 across the 4 of their papers we have counts for
Showing 2019 · cs.CVShow all
2 papers · 2 filters
cs.CV2019
Large-scale representation learning from visually grounded untranscribed speech
Gabriel Ilharco, Yuan Zhang, Jason Baldridge
Systems that can associate images with their spoken audio captions are an important step towards visually grounded language learning. We describe a scalable method to automatically…
cs.CV2019
Transferable Representation Learning in Vision-and-Language Navigation
Haoshuo Huang, Vihan Jain, Harsh Mehta +4
Vision-and-Language Navigation (VLN) tasks such as Room-to-Room (R2R) require machine agents to interpret natural language instructions and learn to act in visually realistic envir…