Adversarial Evaluation of Dialogue Models
arXiv:1701.08198
Abstract
The recent application of RNN encoder-decoder models has resulted in substantial progress in fully data-driven dialogue systems, but evaluation remains a challenge. An adversarial loss could be a way to directly evaluate the extent to which generated dialogue responses sound like they came from a human. This could reduce the need for human evaluation, while more directly evaluating on a generative task. In this work, we investigate this idea by training an RNN to discriminate a dialogue model's samples from human-generated samples. Although we find some evidence this setup could be viable, we also note that many issues remain in its practical application. We discuss both aspects and conclude that future work is warranted.
Cited by in corpus (24)
- Relevance of Unsupervised Metrics in Task-Oriented Dialogue for Evaluating Natural Language Generation
- Challenges in Data-to-Document Generation
- Unifying Human and Statistical Evaluation for Natural Language Generation
- Neural Language Generation: Formulation, Methods, and Evaluation
- Data Distillation for Controlling Specificity in Dialogue Generation
- Designing for Health Chatbots
- PONE: A Novel Automatic Evaluation Metric for Open-Domain Generative Dialogue Systems
- Better Automatic Evaluation of Open-Domain Dialogue Systems with Contextualized Embeddings
- Modeling Text-visual Mutual Dependency for Multi-modal Dialog Generation
- Non-Autoregressive Neural Dialogue Generation
- A Survey of Document Grounded Dialogue Systems (DGDS)
- Logical Natural Language Generation from Open-Domain Tables
- A Measure for Dialog Complexity and its Application in Streamlining Service Operations
- Improved Sentiment Detection via Label Transfer from Monolingual to Synthetic Code-Switched Text
- Teaching Machines to Converse
- Communication-based Evaluation for Natural Language Generation
- How to Evaluate the Next System: Automatic Dialogue Evaluation from the Perspective of Continual Learning
- Learning to Compare for Better Training and Evaluation of Open Domain Natural Language Generation Models
- OpenMEVA: A Benchmark for Evaluating Open-ended Story Generation Metrics
- UNION: An Unreferenced Metric for Evaluating Open-ended Story Generation
- Fine-grained Sentiment Controlled Text Generation
- How to Evaluate Your Dialogue Models: A Review of Approaches
- On the Use of Linguistic Features for the Evaluation of Generative Dialogue Systems
- Predictive Engagement: An Efficient Metric For Automatic Evaluation of Open-Domain Dialogue Systems