315 citations · 955 across the 38 of their papers we have counts for
13 papers · 1 filter
nocaps: novel object captioning at scale
Harsh Agrawal, Karan Desai, Yufei Wang +7
Image captioning models have achieved impressive results on datasets containing limited visual concepts and large amounts of paired image-caption training data. However, if these m…
Do Explanations make VQA Models more Predictable to a Human?
Arjun Chandrasekaran, Viraj Prabhu, Deshraj Yadav +2
A rich line of research attempts to make deep neural networks more transparent by generating human-interpretable 'explanations' of their decision process, especially for interactiv…
TarMAC: Targeted Multi-Agent Communication
Abhishek Das, Théophile Gervet, Joshua Romoff +4
We propose a targeted communication architecture for multi-agent reinforcement learning, where agents learn both what messages to send and whom to address them to while performing…
Neural Modular Control for Embodied Question Answering
Abhishek Das, Georgia Gkioxari, Stefan Lee +2
We present a modular approach for learning policies for navigation over long planning horizons from language input. Our hierarchical policy operates at multiple timescales, where t…
Visual Curiosity: Learning to Ask Questions to Learn Visual Recognition
Jianwei Yang, Jiasen Lu, Stefan Lee +2
In an open-world setting, it is inevitable that an intelligent agent (e.g., a robot) will encounter visual objects, attributes or relationships it does not recognize. In this work,…
Visual Coreference Resolution in Visual Dialog using Neural Module Networks
Satwik Kottur, José M. F. Moura, Devi Parikh +2
Visual dialog entails answering a series of questions grounded in an image, using dialog history as context. In addition to the challenges found in visual question answering (VQA),…