deltaBLEU: A Discriminative Metric for Generation Tasks with Intrinsically Diverse Targets
arXiv:1506.06863
Abstract
We introduce Discriminative BLEU (deltaBLEU), a novel metric for intrinsic evaluation of generated text in tasks that admit a diverse range of possible outputs. Reference strings are scored for quality by human raters on a scale of [-1, +1] to weight multi-reference BLEU. In tasks involving generation of conversational responses, deltaBLEU correlates reasonably with human judgments and outperforms sentence-level and IBM BLEU in terms of both Spearman's rho and Kendall's tau.
6 pages, to appear at ACL 2015
Cited by in corpus (12)
- Relevance of Unsupervised Metrics in Task-Oriented Dialogue for Evaluating Natural Language Generation
- Multi-Task Learning for Speaker-Role Adaptation in Neural Conversation Models
- Topic-based Evaluation for Conversational Bots
- Unifying Human and Statistical Evaluation for Natural Language Generation
- Neural Language Generation: Formulation, Methods, and Evaluation
- Reinforcement Learning Based Emotional Editing Constraint Conversation Generation
- Learning to Start for Sequence to Sequence Architecture
- Report from the NSF Future Directions Workshop, Toward User-Oriented Agents: Research Directions and Challenges
- Language Model Augmented Relevance Score
- Learning to Disambiguate by Asking Discriminative Questions
- How to Evaluate Your Dialogue Models: A Review of Approaches
- On the Use of Linguistic Features for the Evaluation of Generative Dialogue Systems