Why We Need New Evaluation Metrics for NLG
arXiv:1707.06875 · doi:10.18653/v1/D17-1237
Abstract
The majority of NLG evaluation relies on automatic metrics, such as BLEU . In this paper, we motivate the need for novel, system- and data-independent automatic evaluation methods: We investigate a wide range of metrics, including state-of-the-art word-based and novel grammar-based ones, and demonstrate that they only weakly reflect human judgements of system outputs as generated by data-driven, end-to-end NLG. We also show that metric performance is data- and system-specific. Nevertheless, our results also suggest that automatic metrics perform reliably at system-level and can support system development by finding cases where a system performs poorly.
accepted to EMNLP 2017
References in corpus (2)
Cited by in corpus (10)
- Emotionally-Aware Chatbots: A Survey
- Neural Language Generation: Formulation, Methods, and Evaluation
- Robust Conversational AI with Grounded Text Generation
- Distributed Structured Actor-Critic Reinforcement Learning for Universal Dialogue Management
- A Multimodal Dialogue System for Conversational Image Editing
- AgentGraph: Towards Universal Dialogue Management with Structured Deep Reinforcement Learning
- Meta Dialogue Policy Learning
- Document-editing Assistants and Model-based Reinforcement Learning as a Path to Conversational AI
- Hierarchical Context Enhanced Multi-Domain Dialogue System for Multi-domain Task Completion
- Addressing Objects and Their Relations: The Conversational Entity Dialogue Model