Better Automatic Evaluation of Open-Domain Dialogue Systems with Contextualized Embeddings
arXiv:1904.10635
Abstract
Despite advances in open-domain dialogue systems, automatic evaluation of such systems is still a challenging problem. Traditional reference-based metrics such as BLEU are ineffective because there could be many valid responses for a given context that share no common words with reference responses. A recent work proposed Referenced metric and Unreferenced metric Blended Evaluation Routine (RUBER) to combine a learning-based metric, which predicts relatedness between a generated response and a given query, with reference-based metric; it showed high correlation with human judgments. In this paper, we explore using contextualized word embeddings to compute more accurate relatedness scores, thus better evaluation metrics. Experiments show that our evaluation metrics outperform RUBER, which is trained on static embeddings.
8 pages, 2 figures, NAACL 2019 Methods for Optimizing and Evaluating Neural Language Generation (NeuralGen workshop)
References in corpus (1)
Cited by in corpus (6)
- UNION: An Unreferenced Metric for Evaluating Open-ended Story Generation
- Assessing Dialogue Systems with Distribution Distances
- OpenMEVA: A Benchmark for Evaluating Open-ended Story Generation Metrics
- Which Kind Is Better in Open-domain Multi-turn Dialog,Hierarchical or Non-hierarchical Models? An Empirical Study
- Identifying Untrustworthy Samples: Data Filtering for Open-domain Dialogues with Bayesian Optimization
- POSSCORE: A Simple Yet Effective Evaluation of Conversational Search with Part of Speech Labelling