How to Evaluate Your Dialogue Models: A Review of Approaches
arXiv:2108.01369
Abstract
Evaluating the quality of a dialogue system is an understudied problem. The recent evolution of evaluation method motivated this survey, in which an explicit and comprehensive analysis of the existing methods is sought. We are first to divide the evaluation methods into three classes, i.e., automatic evaluation, human-involved evaluation and user simulator based evaluation. Then, each class is covered with main features and the related evaluation metrics. The existence of benchmarks, suitable for the evaluation of dialogue techniques are also discussed in detail. Finally, some open issues are pointed out to bring the evaluation method into a new frontier.
References in corpus (12)
- Relevance of Unsupervised Metrics in Task-Oriented Dialogue for Evaluating Natural Language Generation
- Why We Need New Evaluation Metrics for NLG
- deltaBLEU: A Discriminative Metric for Generation Tasks with Intrinsically Diverse Targets
- Adversarial Evaluation of Dialogue Models
- The GEM Benchmark: Natural Language Generation, its Evaluation and Metrics
- Topic-based Evaluation for Conversational Bots
- Unifying Human and Statistical Evaluation for Natural Language Generation
- ConvLab: Multi-Domain End-to-End Dialog System Platform
- What makes a good conversation? How controllable attributes affect human judgments
- Plato Dialogue System: A Flexible Conversational AI Research Platform
- Neural Multi-task Learning in Automated Assessment
- Re-evaluating ADEM: A Deeper Look at Scoring Dialogue Responses