20 citations · 110 across the 11 of their papers we have counts for
27 papers
Underspecification in Scene Description-to-Depiction Tasks
Ben Hutchinson, Jason Baldridge, Vinodkumar Prabhakaran
Questions regarding implicitness, ambiguity and underspecification are crucial for understanding the task validity and ethical concerns of multimodal image+text systems, yet have r…
MURAL: Multimodal, Multitask Retrieval Across Languages
Aashi Jain, Mandy Guo, Krishna Srinivasan +5
Both image-caption pairs and translation pairs provide the means to learn deep representations of and connections between languages. We use both types of pairs in MURAL (MUltimodal…
Pathdreamer: A World Model for Indoor Navigation
Jing Yu Koh, Honglak Lee, Yinfei Yang +2
People navigating in unfamiliar buildings take advantage of myriad visual, spatial and semantic cues to efficiently achieve their navigation goals. Towards equipping computational…
Talk, Don't Write: A Study of Direct Speech-Based Image Retrieval
Ramon Sanabria, Austin Waters, Jason Baldridge
Speech-based image retrieval has been studied as a proxy for joint representation learning, usually without emphasis on retrieval itself. As such, it is unclear how well speech-bas…
PanGEA: The Panoramic Graph Environment Annotation Toolkit
Alexander Ku, Peter Anderson, Jordi Pont-Tuset +1
PanGEA, the Panoramic Graph Environment Annotation toolkit, is a lightweight toolkit for collecting speech and text annotations in photo-realistic 3D environments. PanGEA immerses…
On the Evaluation of Vision-and-Language Navigation Instructions
Ming Zhao, Peter Anderson, Vihan Jain +4
Vision-and-Language Navigation wayfinding agents can be enhanced by exploiting automatically generated navigation instructions. However, existing instruction generators have not be…