OpenViDial 2.0: A Larger-Scale, Open-Domain Dialogue Generation Dataset with Visual Contexts
arXiv:2109.12761
Abstract
In order to better simulate the real human conversation process, models need to generate dialogue utterances based on not only preceding textual contexts but also visual contexts. However, with the development of multi-modal dialogue learning, the dataset scale gradually becomes a bottleneck. In this report, we release OpenViDial 2.0, a larger-scale open-domain multi-modal dialogue dataset compared to the previous version OpenViDial 1.0. OpenViDial 2.0 contains a total number of 5.6 million dialogue turns extracted from either movies or TV series from different resources, and each dialogue turn is paired with its corresponding visual context. We hope this large-scale dataset can help facilitate future researches on open-domain multi-modal dialog generation, e.g., multi-modal pretraining for dialogue generation.
References in corpus (12)
- Chameleons in imagined conversations: A new approach to understanding coordination of linguistic style in dialogs
- Towards a Human-like Open-Domain Chatbot
- Data Noising as Smoothing in Neural Network Language Models
- Learning Cooperative Visual Dialog Agents with Deep Reinforcement Learning
- A Neural Network Approach to Context-Sensitive Generation of Conversational Responses
- ConvLab: Multi-Domain End-to-End Dialog System Platform
- Rethinking Action Spaces for Reinforcement Learning in End-to-end Dialog Agents with Latent Variable Models
- Explaining Black Box Predictions and Unveiling Data Artifacts through Influence Functions
- Modeling Text-visual Mutual Dependency for Multi-modal Dialog Generation
- Making History Matter: History-Advantage Sequence Training for Visual Dialog
- Non-Autoregressive Neural Dialogue Generation
- Teaching Machines to Converse