RADDLE: An Evaluation Benchmark and Analysis Platform for Robust Task-oriented Dialog Systems
arXiv:2012.14666
Abstract
For task-oriented dialog systems to be maximally useful, it must be able to process conversations in a way that is (1) generalizable with a small number of training examples for new task domains, and (2) robust to user input in various styles, modalities or domains. In pursuit of these goals, we introduce the RADDLE benchmark, a collection of corpora and tools for evaluating the performance of models across a diverse set of domains. By including tasks with limited training data, RADDLE is designed to favor and encourage models with a strong generalization ability. RADDLE also includes a diagnostic checklist that facilitates detailed robustness analysis in aspects such as language variations, speech errors, unseen entities, and out-of-domain utterances. We evaluate recent state-of-the-art systems based on pre-training and fine-tuning, and find that grounded pre-training on heterogeneous dialog corpora performs better than training a separate model per domain. Overall, existing models are less than satisfactory in robustness evaluation, which suggests opportunities for future improvement.
12 pages; Project website at aka.ms/raddle
References in corpus (10)
- Building a Conversational Agent Overnight with Dialogue Self-Play
- DialoGLUE: A Natural Language Understanding Benchmark for Task-Oriented Dialogue
- Adversarial Training for Large Neural Language Models
- A Neural Network Approach to Context-Sensitive Generation of Conversational Responses
- The Eighth Dialog System Technology Challenge
- Few-shot Natural Language Generation for Task-Oriented Dialog
- Task-Oriented Dialog Systems that Consider Multiple Appropriate Responses under the Same Context
- Robust Conversational AI with Grounded Text Generation
- Contextual Out-of-Domain Utterance Handling With Counterfeit Data Augmentation
- Boosting Naturalness of Language in Task-oriented Dialogues via Adversarial Training