Image-Grounded Conversations: Multimodal Context for Natural Question and Response Generation
arXiv:1701.08251
Abstract
The popularity of image sharing on social media and the engagement it creates between users reflects the important role that visual context plays in everyday conversations. We present a novel task, Image-Grounded Conversations (IGC), in which natural-sounding conversations are generated about a shared image. To benchmark progress, we introduce a new multiple-reference dataset of crowd-sourced, event-centric conversations on images. IGC falls on the continuum between chit-chat and goal-directed conversation models, where visual grounding constrains the topic of conversation to event-driven utterances. Experiments with models trained on social media data show that the combination of visual and textual context enhances the quality of generated conversational turns. In human evaluation, the gap between human performance and that of both neural and retrieval architectures suggests that multi-modal IGC presents an interesting challenge for dialogue research.
References in corpus (5)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Sequence to Sequence Learning with Neural Networks
- Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
- Neural Responding Machine for Short-Text Conversation
- deltaBLEU: A Discriminative Metric for Generation Tasks with Intrinsically Diverse Targets
Cited by in corpus (29)
- Best of Both Worlds: Transferring Knowledge from Discriminative Learning to a Generative Visual Dialog Model
- Trends in Integration of Vision and Language Research: A Survey of Tasks, Datasets, and Methods
- C-VQA: A Compositional Split of the Visual Question Answering (VQA) v1.0 Dataset
- Challenges in Building Intelligent Open-domain Dialog Systems
- Multi-step Reasoning via Recurrent Dual Attention for Visual Dialog
- Persona-Based Conversational AI: State of the Art and Challenges
- Active Learning for Visual Question Answering: An Empirical Study
- From FiLM to Video: Multi-turn Question Answering with Multi-modal Context
- Are You Talking to Me? Reasoned Visual Dialog Generation through Adversarial Learning
- Multimodal Hierarchical Reinforcement Learning Policy for Task-Oriented Visual Dialog
- Factor Graph Attention
- CoDraw: Collaborative Drawing as a Testbed for Grounded Goal-driven Communication
- Modeling Text-visual Mutual Dependency for Multi-modal Dialog Generation
- "My Way of Telling a Story": Persona based Grounded Story Generation
- Sequential Attention GAN for Interactive Image Editing
- A System for Automated Image Editing from Natural Language Commands
- Knowledge-Grounded Dialogue Generation with Pre-trained Language Models
- C3VQG: Category Consistent Cyclic Visual Question Generation
- Teaching Machines to Converse
- Learning to Disambiguate by Asking Discriminative Questions
- ZRIGF: An Innovative Multimodal Framework for Zero-Resource Image-Grounded Dialogue Generation
- Constructing Multi-Modal Dialogue Dataset by Replacing Text with Semantically Relevant Images
- Context Retrieval via Normalized Contextual Latent Interaction for Conversational Agent
- Visual Dialogue State Tracking for Question Generation
- Ask No More: Deciding when to guess in referential visual dialogue
- Persona-Coded Poly-Encoder: Persona-Guided Multi-Stream Conversational Sentence Scoring
- Guessing State Tracking for Visual Dialogue
- Game-Based Video-Context Dialogue
- Chat-crowd: A Dialog-based Platform for Visual Layout Composition