How NOT To Evaluate Your Dialogue System: An Empirical Study of Unsupervised Evaluation Metrics for Dialogue Response Generation
arXiv:1603.08023
Abstract
We investigate evaluation metrics for dialogue response generation systems where supervised labels, such as task completion, are not available. Recent works in response generation have adopted metrics from machine translation to compare a model's generated response to a single target response. We show that these metrics correlate very weakly with human judgements in the non-technical Twitter domain, and not at all in the technical Ubuntu domain. We provide quantitative and qualitative results highlighting specific weaknesses in existing metrics, and provide recommendations for future development of better automatic evaluation metrics for dialogue systems.
First 4 authors had equal contribution. 13 pages, 5 tables, 6 figures. EMNLP 2016
References in corpus (2)
Cited by in corpus (230)
- A Deep Reinforced Model for Abstractive Summarization
- DailyDialog: A Manually Labelled Multi-turn Dialogue Dataset
- Deep Reinforcement Learning for Dialogue Generation
- TransferTransfo: A Transfer Learning Approach for Neural Network Based Conversational Agents
- A Simple, Fast Diverse Decoding Algorithm for Neural Generation
- Adversarial Learning for Neural Dialogue Generation
- A Deep Reinforcement Learning Chatbot
- Evaluation of Text Generation: A Survey
- Personalizing Dialogue Agents: I have a dog, do you have pets too?
- Relevance of Unsupervised Metrics in Task-Oriented Dialogue for Evaluating Natural Language Generation
- Cyclical Annealing Schedule: A Simple Approach to Mitigating KL Vanishing
- Learning Discourse-level Diversity for Neural Dialog Models using Conditional Variational Autoencoders
- Clinically Accurate Chest X-Ray Report Generation
- A Hierarchical Structured Self-Attentive Model for Extractive Document Summarization (HSSAS)
- DialogWAE: Multimodal Response Generation with Conditional Wasserstein Auto-Encoder
- Emotional Chatting Machine: Emotional Conversation Generation with Internal and External Memory
- CoaCor: Code Annotation for Code Retrieval with Reinforcement Learning
- Recipes for Safety in Open-domain Chatbots
- An End-to-End Trainable Neural Network Model with Belief Tracking for Task-Oriented Dialog
- Two are Better than One: An Ensemble of Retrieval- and Generation-Based Dialog Systems
- Neural Models for Information Retrieval
- Best of Both Worlds: Transferring Knowledge from Discriminative Learning to a Generative Visual Dialog Model
- Recent Advances in Neural Question Generation
- Language GANs Falling Short
- NewsQA: A Machine Comprehension Dataset
- Learning End-to-End Goal-Oriented Dialog
- Chinese Poetry Generation with Planning based Neural Network
- Adversarial Evaluation of Dialogue Models
- Multi-Task Learning for Speaker-Role Adaptation in Neural Conversation Models
- DP-GAN: Diversity-Promoting Generative Adversarial Network for Generating Informative and Diversified Text
- Challenges in Data-to-Document Generation
- Zero-Resource Knowledge-Grounded Dialogue Generation
- Plan-And-Write: Towards Better Automatic Storytelling
- Variational Transformers for Diverse Response Generation
- Semi-Supervised Variational Reasoning for Medical Dialogue Generation
- Meta-evaluation of Conversational Search Evaluation Metrics
- EvalAI: Towards Better Evaluation Systems for AI Agents
- PaperRobot: Incremental Draft Generation of Scientific Ideas
- The price of debiasing automatic metrics in natural language evaluation
- Topic-based Evaluation for Conversational Bots
- Unifying Human and Statistical Evaluation for Natural Language Generation
- Deep Learning Based Chatbot Models
- Keyphrase Extraction from Disaster-related Tweets
- Emergence of Compositional Language with Deep Generational Transmission
- Learning Symmetric Collaborative Dialogue Agents with Dynamic Knowledge Graph Embeddings
- Neural Text Generation: A Practical Guide
- Asking and Answering Questions to Evaluate the Factual Consistency of Summaries
- Improving Variational Encoder-Decoders in Dialogue Generation
- End-to-end optimization of goal-driven and visually grounded dialogue systems
- Image Inspired Poetry Generation in XiaoIce
- Towards Explainable and Controllable Open Domain Dialogue Generation with Dialogue Acts
- Latent Variable Dialogue Models and their Diversity
- A Conditional Variational Framework for Dialog Generation
- EVA: An Open-Domain Chinese Dialogue System with Large-Scale Generative Pre-Training
- What makes a good conversation? How controllable attributes affect human judgments
- An Attentional Neural Conversation Model with Improved Specificity
- Natural Language Generation in Dialogue using Lexicalized and Delexicalized Data
- On Evaluating and Comparing Open Domain Dialog Systems
- End-to-end Conversation Modeling Track in DSTC6
- Sequence Tutor: Conservative Fine-Tuning of Sequence Generation Models with KL-control
- No Metrics Are Perfect: Adversarial Reward Learning for Visual Storytelling
- Neural Language Generation: Formulation, Methods, and Evaluation
- CPM: A Large-scale Generative Chinese Pre-trained Language Model
- An Adversarial Approach to High-Quality, Sentiment-Controlled Neural Dialogue Generation
- MoverScore: Text Generation Evaluating with Contextualized Embeddings and Earth Mover Distance
- Variational Attention for Sequence-to-Sequence Models
- GuessWhat?! Visual object discovery through multi-modal dialogue
- Domain Aware Neural Dialog System
- Generating High-Quality and Informative Conversation Responses with Sequence-to-Sequence Models
- Data Distillation for Controlling Specificity in Dialogue Generation
- Bidirectional Attentional Encoder-Decoder Model and Bidirectional Beam Search for Abstractive Summarization
- From FiLM to Video: Multi-turn Question Answering with Multi-modal Context
- Exploiting Persona Information for Diverse Generation of Conversational Responses
- Designing for Health Chatbots
- Retrieve and Refine: Improved Sequence Generation Models For Dialogue
- PLATO-XL: Exploring the Large-scale Pre-training of Dialogue Generation
- Production Ready Chatbots: Generate if not Retrieve
- A Hierarchical Latent Structure for Variational Conversation Modeling
- A Neural Topical Expansion Framework for Unstructured Persona-oriented Dialogue Generation
- Better Automatic Evaluation of Open-Domain Dialogue Systems with Contextualized Embeddings
- Automatic Evaluation and Moderation of Open-domain Dialogue Systems
- Referenceless Quality Estimation for Natural Language Generation
- PLATO: Pre-trained Dialogue Generation Model with Discrete Latent Variable
- Towards Coherent and Engaging Spoken Dialog Response Generation Using Automatic Conversation Evaluators
- Alquist: The Alexa Prize Socialbot
- RubyStar: A Non-Task-Oriented Mixture Model Dialog System
- CDL: Curriculum Dual Learning for Emotion-Controllable Response Generation
- Steering Output Style and Topic in Neural Response Generation
- Automatic Article Commenting: the Task and Dataset
- Non-Autoregressive Neural Dialogue Generation
- Beyond Turing: Intelligent Agents Centered on the User
- On the Generation of Medical Dialogues for COVID-19
- Going Beneath the Surface: Evaluating Image Captioning for Grammaticality, Truthfulness and Diversity
- Answers Unite! Unsupervised Metrics for Reinforced Summarization Models
- Chat More If You Like: Dynamic Cue Words Planning to Flow Longer Conversations
- Natural Language Generation with Neural Variational Models
- Adaptive Parameterization for Neural Dialogue Generation
- Are Pre-trained Language Models Knowledgeable to Ground Open Domain Dialogues?
- End-to-end Adversarial Learning for Generative Conversational Agents
- An Entity-Driven Framework for Abstractive Summarization
- CoMAE: A Multi-factor Hierarchical Framework for Empathetic Response Generation
- Topic-Preserving Synthetic News Generation: An Adversarial Deep Reinforcement Learning Approach
- Towards Emotional Support Dialog Systems
- Probing Neural Dialog Models for Conversational Understanding
- Improving Conditional Sequence Generative Adversarial Networks by Stepwise Evaluation
- DynaEval: Unifying Turn and Dialogue Level Evaluation
- MoEL: Mixture of Empathetic Listeners
- Towards a Metric for Automated Conversational Dialogue System Evaluation and Improvement
- GRADE: Automatic Graph-Enhanced Coherence Metric for Evaluating Open-Domain Dialogue Systems
- End-to-End Knowledge-Routed Relational Dialogue System for Automatic Diagnosis
- A Syntactically Constrained Bidirectional-Asynchronous Approach for Emotional Conversation Generation
- Generating Multiple Diverse Responses with Multi-Mapping and Posterior Mapping Selection
- Report from the NSF Future Directions Workshop, Toward User-Oriented Agents: Research Directions and Challenges
- Action-Based Conversations Dataset: A Corpus for Building More In-Depth Task-Oriented Dialogue Systems
- Ultra-Fast, Low-Storage, Highly Effective Coarse-grained Selection in Retrieval-based Chatbot by Using Deep Semantic Hashing
- Cue Me In: Content-Inducing Approaches to Interactive Story Generation
- Differentiable lower bound for expected BLEU score
- The Rapidly Changing Landscape of Conversational Agents
- Learning from Fact-checkers: Analysis and Generation of Fact-checking Language
- Communication-based Evaluation for Natural Language Generation
- The JDDC 2.0 Corpus: A Large-Scale Multimodal Multi-Turn Chinese Dialogue Dataset for E-commerce Customer Service
- Generating More Interesting Responses in Neural Conversation Models with Distributional Constraints
- How to Evaluate the Next System: Automatic Dialogue Evaluation from the Perspective of Continual Learning
- CRSLab: An Open-Source Toolkit for Building Conversational Recommender System
- Polite Dialogue Generation Without Parallel Data
- Knowledge-Grounded Dialogue Generation with Pre-trained Language Models
- Implicit Deep Latent Variable Models for Text Generation
- ARAML: A Stable Adversarial Training Framework for Text Generation
- Counterfactual Story Reasoning and Generation
- Retrieval Augmentation Reduces Hallucination in Conversation
- A Large-Scale Chinese Short-Text Conversation Dataset
- Constructing Emotion Consensus and Utilizing Unpaired Data for Empathetic Dialogue Generation
- Judge the Judges: A Large-Scale Evaluation Study of Neural Language Models for Online Review Generation
- Negative Training for Neural Dialogue Response Generation
- How To Evaluate Your Dialogue System: Probe Tasks as an Alternative for Token-level Evaluation Metrics
- Improving Question Answering Model Robustness with Synthetic Adversarial Data Generation
- Towards a Better Metric for Evaluating Question Generation Systems
- Unsupervised Learning of KB Queries in Task-Oriented Dialogs
- I Do Not Understand What I Cannot Define: Automatic Question Generation With Pedagogically-Driven Content Selection
- Latent Topic Conversational Models
- Designing dialogue systems: A mean, grumpy, sarcastic chatbot in the browser
- Being Negative but Constructively: Lessons Learnt from Creating Better Visual Question Answering Datasets
- An Evaluation Protocol for Generative Conversational Systems
- Fine-Grained Sentence Functions for Short-Text Conversation
- Learning to Compare for Better Training and Evaluation of Open Domain Natural Language Generation Models
- APo-VAE: Text Generation in Hyperbolic Space
- Affective Neural Response Generation
- Generating Personalized Dialogue via Multi-Task Meta-Learning
- Modeling Semantic Relationship in Multi-turn Conversations with Hierarchical Latent Variables
- A Discrete CVAE for Response Generation on Short-Text Conversation
- UNION: An Unreferenced Metric for Evaluating Open-ended Story Generation
- How to Evaluate Your Dialogue Models: A Review of Approaches
- OpenMEVA: A Benchmark for Evaluating Open-ended Story Generation Metrics
- What is More Likely to Happen Next? Video-and-Language Future Event Prediction
- How to Build User Simulators to Train RL-based Dialog Systems
- Aiming to Know You Better Perhaps Makes Me a More Engaging Dialogue Partner
- Latent Semantic Analysis Approach for Document Summarization Based on Word Embeddings
- Modeling Event Background for If-Then Commonsense Reasoning Using Context-aware Variational Autoencoder
- Tree-Structured Neural Machine for Linguistics-Aware Sentence Generation
- Generating Dialogue Responses from a Semantic Latent Space
- Dialogue Generation on Infrequent Sentence Functions via Structured Meta-Learning
- Explain Me the Painting: Multi-Topic Knowledgeable Art Description Generation
- Evaluating Dialogue Generation Systems via Response Selection
- EnsembleGAN: Adversarial Learning for Retrieval-Generation Ensemble Model on Short-Text Conversation
- One "Ruler" for All Languages: Multi-Lingual Dialogue Evaluation with Adversarial Multi-Task Learning
- Contextual Topic Modeling For Dialog Systems
- UniConv: A Unified Conversational Neural Architecture for Multi-domain Task-oriented Dialogues
- A Natural Language Corpus of Common Grounding under Continuous and Partially-Observable Context
- Directed Beam Search: Plug-and-Play Lexically Constrained Language Generation
- Generating Chinese Poetry from Images via Concrete and Abstract Information
- Guiding Variational Response Generator to Exploit Persona
- Combining Textual Content and Structure to Improve Dialog Similarity
- Cue-word Driven Neural Response Generation with a Shrinking Vocabulary
- LEGOEval: An Open-Source Toolkit for Dialogue System Evaluation via Crowdsourcing
- Semantic-Enhanced Explainable Finetuning for Open-Domain Dialogues
- Know More about Each Other: Evolving Dialogue Strategy via Compound Assessment
- Filtering Noisy Dialogue Corpora by Connectivity and Content Relatedness
- Increasing Faithfulness in Knowledge-Grounded Dialogue with Controllable Features
- A Multi-Turn Emotionally Engaging Dialog Model
- Plan ahead: Self-Supervised Text Planning for Paragraph Completion Task
- Commonsense-Focused Dialogues for Response Generation: An Empirical Study
- EmpBot: A T5-based Empathetic Chatbot focusing on Sentiments
- Counterfactual Off-Policy Training for Neural Response Generation
- Achieving Fluency and Coherency in Task-oriented Dialog
- Local Knowledge Powered Conversational Agents
- On the Use of Linguistic Features for the Evaluation of Generative Dialogue Systems
- Learning to Predict Persona Information forDialogue Personalization without Explicit Persona Description
- Multi-Task Learning of Generation and Classification for Emotion-Aware Dialogue Response Generation
- Which Kind Is Better in Open-domain Multi-turn Dialog,Hierarchical or Non-hierarchical Models? An Empirical Study
- Opinion-aware Answer Generation for Review-driven Question Answering in E-Commerce
- DyKgChat: Benchmarking Dialogue Generation Grounding on Dynamic Knowledge Graphs
- : Evaluating Factual Consistency in Knowledge-Grounded Dialogues via Question Generation and Question Answering
- Improving Computer Generated Dialog with Auxiliary Loss Functions and Custom Evaluation Metrics
- Posterior-GAN: Towards Informative and Coherent Response Generation with Posterior Generative Adversarial Network
- Review-Driven Answer Generation for Product-Related Questions in E-Commerce
- A Corpus of Controlled Opinionated and Knowledgeable Movie Discussions for Training Neural Conversation Models
- Pchatbot: A Large-Scale Dataset for Personalized Chatbot
- Conversational Response Re-ranking Based on Event Causality and Role Factored Tensor Event Embedding
- Positioning yourself in the maze of Neural Text Generation: A Task-Agnostic Survey
- (Male, Bachelor) and (Female, Ph.D) have different connotations: Parallelly Annotated Stylistic Language Dataset with Multiple Personas
- Linguistic Versus Latent Relations for Modeling Coherent Flow in Paragraphs
- Predicting User Engagement Status for Online Evaluation of Intelligent Assistants
- Multichannel Generative Language Model: Learning All Possible Factorizations Within and Across Channels
- Multiple Generative Models Ensemble for Knowledge-Driven Proactive Human-Computer Dialogue Agent
- Towards Automatic Evaluation of Dialog Systems: A Model-Free Off-Policy Evaluation Approach
- REAM: An Enhancement Approach to Reference-based Evaluation Metrics for Open-domain Dialog Generation
- Semantic Extractor-Paraphraser based Abstractive Summarization
- The RLLChatbot: a solution to the ConvAI challenge
- User Response and Sentiment Prediction for Automatic Dialogue Evaluation
- CO-STAR: Conceptualisation of Stereotypes for Analysis and Reasoning
- Automatic Evaluation of Neural Personality-based Chatbots
- Learning from Perturbations: Diverse and Informative Dialogue Generation with Inverse Adversarial Training
- A Dataset and Baselines for Multilingual Reply Suggestion
- GTM: A Generative Triple-Wise Model for Conversational Question Generation
- Generating Relevant and Coherent Dialogue Responses using Self-separated Conditional Variational AutoEncoders
- Do Encoder Representations of Generative Dialogue Models Encode Sufficient Information about the Task ?
- Towards Quantifiable Dialogue Coherence Evaluation
- Building Chatbots from Forum Data: Model Selection Using Question Answering Metrics
- POSSCORE: A Simple Yet Effective Evaluation of Conversational Search with Part of Speech Labelling
- Conversational Multi-Hop Reasoning with Neural Commonsense Knowledge and Symbolic Logic Rules
- DiSCoL: Toward Engaging Dialogue Systems through Conversational Line Guided Response Generation
- An Emotion-controlled Dialog Response Generation Model with Dynamic Vocabulary
- Estimating Subjective Crowd-Evaluations as an Additional Objective to Improve Natural Language Generation
- Speaker Sensitive Response Evaluation Model
- Evidence-Aware Inferential Text Generation with Vector Quantised Variational AutoEncoder
- Focus-Constrained Attention Mechanism for CVAE-based Response Generation
- Unsupervised Context Rewriting for Open Domain Conversation
- Classification as Decoder: Trading Flexibility for Control in Medical Dialogue
- Personalizing Dialogue Agents via Meta-Learning
- Target-Guided Open-Domain Conversation