315 citations · 955 across the 39 of their papers we have counts for
7 papers · 2 filters
Large-scale Pretraining for Visual Dialog: A Simple State-of-the-Art Baseline
Vishvak Murahari, Dhruv Batra, Devi Parikh +1
Prior work in visual dialog has focused on training deep neural models on VisDial in isolation. Instead, we present an approach to leverage pretraining on related vision-language d…
Improving Generative Visual Dialog by Answering Diverse Questions
Vishvak Murahari, Prithvijit Chattopadhyay, Dhruv Batra +2
Prior work on training generative Visual Dialog models with reinforcement learning(Das et al.) has explored a Qbot-Abot image-guessing game and shown that this 'self-talk' approach…
IR-VIC: Unsupervised Discovery of Sub-goals for Transfer in RL
Nirbhay Modhe, Prithvijit Chattopadhyay, Mohit Sharma +4
We propose a novel framework to identify sub-goals useful for exploration in sequential decision making tasks under partial observability. We utilize the variational intrinsic cont…
Emergence of Compositional Language with Deep Generational Transmission
Michael Cogswell, Jiasen Lu, Stefan Lee +2
Recent work has studied the emergence of language among deep reinforcement learning agents that must collaborate to solve a task. Of particular interest are the factors that cause…
Counterfactual Visual Explanations
Yash Goyal, Ziyan Wu, Jan Ernst +3
In this work, we develop a technique to produce counterfactual visual explanations. Given a 'query' image for which a vision system predicts class , a counterfactual visual…
Embodied Multimodal Multitask Learning
Devendra Singh Chaplot, Lisa Lee, Ruslan Salakhutdinov +2
Recent efforts on training visual navigation agents conditioned on language using deep reinforcement learning have been successful in learning policies for different multimodal tas…