5 papers
Episodic Memory Question Answering
Samyak Datta, Sameer Dharur, Vincent Cartillier +4
Egocentric augmented reality devices such as wearable glasses passively capture visual data as a human wearer tours a home environment. We envision a scenario wherein the human com…
Semantic MapNet: Building Allocentric Semantic Maps and Representations from Egocentric Views
Vincent Cartillier, Zhile Ren, Neha Jain +3
We study the task of semantic mapping - specifically, an embodied agent (a robot or an egocentric AI assistant) is given a tour of a new environment and asked to build an allocentr…
Audio-Visual Scene-Aware Dialog
Huda Alamri, Vincent Cartillier, Abhishek Das +9
We introduce the task of scene-aware dialog. Our goal is to generate a complete and natural response to a question about a scene, given video and audio of the scene and the history…
End-to-End Audio Visual Scene-Aware Dialog using Multimodal Attention-Based Video Features
Chiori Hori, Huda Alamri, Jue Wang +10
Dialog systems need to understand dynamic visual scenes in order to have conversations with users about the objects and events around them. Scene-aware dialog systems for real-worl…
Audio Visual Scene-Aware Dialog (AVSD) Challenge at DSTC7
Huda Alamri, Vincent Cartillier, Raphael Gontijo Lopes +8
Scene-aware dialog systems will be able to have conversations with users about the objects and events around them. Progress on such systems can be made by integrating state-of-the-…