37 citations · 129 across the 10 of their papers we have counts for
10 papers
FineD-Eval: Fine-grained Automatic Dialogue-Level Evaluation
Chen Zhang, Luis Fernando D'Haro, Qiquan Zhang +2
Recent model-based reference-free metrics for open-domain dialogue evaluation exhibit promising correlations with human judgment. However, they either perform turn-level evaluation…
Report from the NSF Future Directions Workshop on Automatic Evaluation of Dialog: Research Directions and Challenges
Shikib Mehri, Jinho Choi, Luis Fernando D'Haro +13
This is a report on the NSF Future Directions Workshop on Automatic Evaluation of Dialog. The workshop explored the current state of the art along with its limitations and suggeste…
MDD-Eval: Self-Training on Augmented Data for Multi-Domain Dialogue Evaluation
Chen Zhang, Luis Fernando D'Haro, Thomas Friedrichs +1
Chatbots are designed to carry out human-like conversations across different domains, such as general chit-chat, knowledge exchange, and persona-grounded conversations. To measure…
Investigating the Impact of Pre-trained Language Models on Dialog Evaluation
Chen Zhang, Luis Fernando D'Haro, Yiming Chen +2
Recently, there is a surge of interest in applying pre-trained language models (Pr-LM) in automatic open-domain dialog evaluation. Pr-LMs offer a promising direction for addressing…
DynaEval: Unifying Turn and Dialogue Level Evaluation
Chen Zhang, Yiming Chen, Luis Fernando D'Haro +4
A dialogue is essentially a multi-turn interaction among interlocutors. Effective evaluation metrics should reflect the dynamics of such interaction. Existing automatic metrics are…
Overview of the Ninth Dialog System Technology Challenge: DSTC9
Chulaka Gunasekara, Seokhwan Kim, Luis Fernando D'Haro +36
This paper introduces the Ninth Dialog System Technology Challenge (DSTC-9). This edition of the DSTC focuses on applying end-to-end dialog technologies for four distinct tasks in…