Towards Standard Criteria for human evaluation of Chatbots: A Survey
arXiv:2105.11197
Abstract
Human evaluation is becoming a necessity to test the performance of Chatbots. However, off-the-shelf settings suffer the severe reliability and replication issues partly because of the extremely high diversity of criteria. It is high time to come up with standard criteria and exact definitions. To this end, we conduct a through investigation of 105 papers involving human evaluation for Chatbots. Deriving from this, we propose five standard criteria along with precise definitions.
References in corpus (6)
- Towards a Human-like Open-Domain Chatbot
- Estimation-Action-Reflection: Towards Deep Interaction Between Conversational and Recommender Systems
- Recent Advances in Neural Question Generation
- What makes a good conversation? How controllable attributes affect human judgments
- Diversifying Dialogue Generation with Non-Conversational Text
- Generating Dialogue Responses from a Semantic Latent Space