Learning Cooperative Visual Dialog Agents with Deep Reinforcement Learning
arXiv:1703.06585
Abstract
We introduce the first goal-driven training for visual question answering and dialog agents. Specifically, we pose a cooperative 'image guessing' game between two agents -- Qbot and Abot -- who communicate in natural language dialog so that Qbot can select an unseen image from a lineup of images. We use deep reinforcement learning (RL) to learn the policies of these agents end-to-end -- from pixels to multi-agent multi-round dialog to game reward. We demonstrate two experimental results. First, as a 'sanity check' demonstration of pure RL (from scratch), we show results on a synthetic world, where the agents communicate in ungrounded vocabulary, i.e., symbols with no pre-specified meanings (X, Y, Z). We find that two bots invent their own communication protocol and start using certain symbols to ask/answer about certain visual attributes (shape/color/style). Thus, we demonstrate the emergence of grounded language and communication among 'visual' dialog agents with no human supervision. Second, we conduct large-scale real-image experiments on the VisDial dataset, where we pretrain with supervised dialog data and show that the RL 'fine-tuned' agents significantly outperform SL agents. Interestingly, the RL Qbot learns to ask questions that Abot is good at, ultimately resulting in more informative dialog and a better team.
11 pages, 4 figures, 2 tables, webpage: http://visualdialog.org/
References in corpus (12)
- Adam: A Method for Stochastic Optimization
- Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
- VQA: Visual Question Answering
- Learning to Communicate with Deep Multi-Agent Reinforcement Learning
- Deep Reinforcement Learning for Dialogue Generation
- Are You Talking to a Machine? Dataset and Methods for Multilingual Image Question Answering
- Adversarial Learning for Neural Dialogue Generation
- Sequence to Sequence -- Video to Text
- Show and Tell: A Neural Image Caption Generator
- Ask Your Neurons: A Neural-based Approach to Answering Questions about Images
- From Captions to Visual Concepts and Back
- GuessWhat?! Visual object discovery through multi-modal dialogue
Cited by in corpus (57)
- Deep Reinforcement Learning: An Overview
- A Deep Reinforcement Learning Chatbot
- Dialog-based Interactive Image Retrieval
- Maintaining cooperation in complex social dilemmas using deep reinforcement learning
- Visual Reference Resolution using Attention Memory for Visual Dialog
- Best of Both Worlds: Transferring Knowledge from Discriminative Learning to a Generative Visual Dialog Model
- Discriminability objective for training descriptive captions
- Modeling Others using Oneself in Multi-Agent Reinforcement Learning
- Out of the Box: Reasoning with Graph Convolution Nets for Factual Visual Question Answering
- Emergence of Language with Multi-agent Games: Learning to Communicate with Sequences of Symbols
- Emergent Translation in Multi-Agent Communication
- C-VQA: A Compositional Split of the Visual Question Answering (VQA) v1.0 Dataset
- EvalAI: Towards Better Evaluation Systems for AI Agents
- Deep Learning Based Chatbot Models
- ACCNet: Actor-Coordinator-Critic Net for "Learning-to-Communicate" with Deep Multi-agent Reinforcement Learning
- Composite Task-Completion Dialogue Policy Learning via Hierarchical Deep Reinforcement Learning
- Emergent Communication in a Multi-Modal, Multi-Step Referential Game
- Audio Visual Scene-Aware Dialog (AVSD) Challenge at DSTC7
- Context-aware Captions from Context-agnostic Supervision
- Federated Control with Hierarchical Multi-Agent Deep Reinforcement Learning
- Controllable Video Captioning with an Exemplar Sentence
- Visual Dialogue without Vision or Dialogue
- Consequentialist conditional cooperation in social dilemmas with imperfect information
- Active Learning for Visual Question Answering: An Empirical Study
- From FiLM to Video: Multi-turn Question Answering with Multi-modal Context
- Asking the Difficult Questions: Goal-Oriented Visual Question Generation via Intermediate Rewards
- Grounding Referring Expressions in Images by Variational Context
- Are You Talking to Me? Reasoned Visual Dialog Generation through Adversarial Learning
- FlipDial: A Generative Model for Two-Way Visual Dialogue
- Multimodal Hierarchical Reinforcement Learning Policy for Task-Oriented Visual Dialog
- CoDraw: Collaborative Drawing as a Testbed for Grounded Goal-driven Communication
- Modeling Text-visual Mutual Dependency for Multi-modal Dialog Generation
- Interpretable and Pedagogical Examples
- End-to-End Audio Visual Scene-Aware Dialog using Multimodal Attention-Based Video Features
- Arena: A General Evaluation Platform and Building Toolkit for Multi-Agent Intelligence
- Learning Autocomplete Systems as a Communication Game
- OpenViDial 2.0: A Larger-Scale, Open-Domain Dialogue Generation Dataset with Visual Contexts
- Beyond task success: A closer look at jointly learning to see, ask, and GuessWhat
- Variational Inference for Data-Efficient Model Learning in POMDPs
- Interactive Language Acquisition with One-shot Visual Concept Learning through a Conversational Game
- Outline Objects using Deep Reinforcement Learning
- Examining Cooperation in Visual Dialog Models
- Bridging the Imitation Gap by Adaptive Insubordination
- Visual Dialogue State Tracking for Question Generation
- The emergence of visual semantics through communication games
- Ask No More: Deciding when to guess in referential visual dialogue
- Explanation of Reinforcement Learning Model in Dynamic Multi-Agent System
- Parallel Attention: A Unified Framework for Visual Object Discovery through Dialogs and Queries
- Deep Reinforcement Learning for Active Human Pose Estimation
- Semantic bottleneck for computer vision tasks
- Multi-modal dialog for browsing large visual catalogs using exploration-exploitation paradigm in a joint embedding space
- Interactive Video Retrieval with Dialog
- Saying the Unseen: Video Descriptions via Dialog Agents
- Guessing State Tracking for Visual Dialogue
- Reasoning Over History: Context Aware Visual Dialog
- Live Video Comment Generation Based on Surrounding Frames and Live Comments
- Emergent Communication with World Models