Relevance of Unsupervised Metrics in Task-Oriented Dialogue for Evaluating Natural Language Generation
arXiv:1706.09799
Abstract
Automated metrics such as BLEU are widely used in the machine translation literature. They have also been used recently in the dialogue community for evaluating dialogue response generation. However, previous work in dialogue response generation has shown that these metrics do not correlate strongly with human judgment in the non task-oriented dialogue setting. Task-oriented dialogue responses are expressed on narrower domains and exhibit lower diversity. It is thus reasonable to think that these automated metrics would correlate well with human judgment in the task-oriented setting where the generation task consists of translating dialogue acts into a sentence. We conduct an empirical study to confirm whether this is the case. Our findings indicate that these automated metrics have stronger correlation with human judgments in the task-oriented setting compared to what has been observed in the non task-oriented setting. We also observe that these metrics correlate even better for datasets which provide multiple ground truth reference sentences. In addition, we show that some of the currently available corpora for task-oriented language generation can be solved with simple models and advocate for more challenging datasets.
References in corpus (3)
Cited by in corpus (13)
- End-to-end Conversation Modeling Track in DSTC6
- From FiLM to Video: Multi-turn Question Answering with Multi-modal Context
- Step-by-Step: Separating Planning from Realization in Neural Data-to-Text Generation
- Contrastive Learning with Adversarial Perturbations for Conditional Text Generation
- Evaluating Style Transfer for Text
- Communication-based Evaluation for Natural Language Generation
- SMRT Chatbots: Improving Non-Task-Oriented Dialog with Simulated Multiple Reference Training
- Knowledge-Grounded Response Generation with Deep Attentional Latent-Variable Model
- Adversarial Black-Box Attacks On Text Classifiers Using Multi-Objective Genetic Optimization Guided By Deep Networks
- A Revised Generative Evaluation of Visual Dialogue
- Data-Efficient Methods for Dialogue Systems
- Diamonds in the Rough: Generating Fluent Sentences from Early-Stage Drafts for Academic Writing Assistance
- Extended Answer and Uncertainty Aware Neural Question Generation