RUBER: An Unsupervised Method for Automatic Evaluation of Open-Domain Dialog Systems
arXiv:1701.03079
Abstract
Open-domain human-computer conversation has been attracting increasing attention over the past few years. However, there does not exist a standard automatic evaluation metric for open-domain dialog systems; researchers usually resort to human annotation for model evaluation, which is time- and labor-intensive. In this paper, we propose RUBER, a Referenced metric and Unreferenced metric Blended Evaluation Routine, which evaluates a reply by taking into consideration both a groundtruth reply and a query (previous user-issued utterance). Our metric is learnable, but its training does not require labels of human satisfaction. Hence, RUBER is flexible and extensible to different datasets and languages. Experiments on both retrieval and generative dialog systems show that RUBER has a high correlation with human annotation.
Cited by in corpus (44)
- Meta-evaluation of Conversational Search Evaluation Metrics
- Challenges in Building Intelligent Open-domain Dialog Systems
- Neural Language Generation: Formulation, Methods, and Evaluation
- Towards Topic-Guided Conversational Recommender System
- PONE: A Novel Automatic Evaluation Metric for Open-Domain Generative Dialogue Systems
- Better Automatic Evaluation of Open-Domain Dialogue Systems with Contextualized Embeddings
- A Survey of Document Grounded Dialogue Systems (DGDS)
- Towards Standard Criteria for human evaluation of Chatbots: A Survey
- Meaningful Answer Generation of E-Commerce Question-Answering
- Beyond User Self-Reported Likert Scale Ratings: A Comparison Model for Automatic Dialog Evaluation
- DynaEval: Unifying Turn and Dialogue Level Evaluation
- Ultra-Fast, Low-Storage, Highly Effective Coarse-grained Selection in Retrieval-based Chatbot by Using Deep Semantic Hashing
- GRADE: Automatic Graph-Enhanced Coherence Metric for Evaluating Open-Domain Dialogue Systems
- Designing Precise and Robust Dialogue Response Evaluators
- ARAML: A Stable Adversarial Training Framework for Text Generation
- How To Evaluate Your Dialogue System: Probe Tasks as an Alternative for Token-level Evaluation Metrics
- Language Model Augmented Relevance Score
- How to Evaluate Your Dialogue Models: A Review of Approaches
- Product-Aware Answer Generation in E-Commerce Question-Answering
- UNION: An Unreferenced Metric for Evaluating Open-ended Story Generation
- Assessing Dialogue Systems with Distribution Distances
- OpenMEVA: A Benchmark for Evaluating Open-ended Story Generation Metrics
- Deconstruct to Reconstruct a Configurable Evaluation Metric for Open-Domain Dialogue Systems
- EnsembleGAN: Adversarial Learning for Retrieval-Generation Ensemble Model on Short-Text Conversation
- One "Ruler" for All Languages: Multi-Lingual Dialogue Evaluation with Adversarial Multi-Task Learning
- The statistical advantage of automatic NLG metrics at the system level
- CLIPScore: A Reference-free Evaluation Metric for Image Captioning
- Commonsense-Focused Dialogues for Response Generation: An Empirical Study
- A Multi-Turn Emotionally Engaging Dialog Model
- Which Kind Is Better in Open-domain Multi-turn Dialog,Hierarchical or Non-hierarchical Models? An Empirical Study
- Filtering Noisy Dialogue Corpora by Connectivity and Content Relatedness
- POSSCORE: A Simple Yet Effective Evaluation of Conversational Search with Part of Speech Labelling
- REAM: An Enhancement Approach to Reference-based Evaluation Metrics for Open-domain Dialog Generation
- Building A User-Centric and Content-Driven Socialbot
- Towards Automatic Evaluation of Dialog Systems: A Model-Free Off-Policy Evaluation Approach
- Speaker Sensitive Response Evaluation Model
- Towards Quantifiable Dialogue Coherence Evaluation
- Do Encoder Representations of Generative Dialogue Models Encode Sufficient Information about the Task ?
- GTM: A Generative Triple-Wise Model for Conversational Question Generation
- HERALD: An Annotation Efficient Method to Detect User Disengagement in Social Conversations
- DiSCoL: Toward Engaging Dialogue Systems through Conversational Line Guided Response Generation
- On conducting better validation studies of automatic metrics in natural language generation evaluation
- Predictive Engagement: An Efficient Metric For Automatic Evaluation of Open-Domain Dialogue Systems
- User Response and Sentiment Prediction for Automatic Dialogue Evaluation