most citedReport from the NSF Future Directions Workshop on Automatic Evaluation of Dialog: Research Directions and Challenges

24 citations · 50 across the 8 of their papers we have counts for

collaborators
Showing cs.CLShow all

8 papers · 1 filter

cs.CL2022

FineD-Eval: Fine-grained Automatic Dialogue-Level Evaluation

Chen Zhang, Luis Fernando D'Haro, Qiquan Zhang +2

Recent model-based reference-free metrics for open-domain dialogue evaluation exhibit promising correlations with human judgment. However, they either perform turn-level evaluation…

cs.CL202210 cited

A Focused Study on Sequence Length for Dialogue Summarization

Bin Wang, Chen Zhang, Chengwei Wei +1

Output length is critical to dialogue summarization systems. The dialogue summary length is determined by multiple factors, including dialogue complexity, summary objective, and pe…

cs.CL2022

Analyzing and Evaluating Faithfulness in Dialogue Summarization

Bin Wang, Chen Zhang, Yan Zhang +2

Dialogue summarization is abstractive in nature, making it suffer from factual errors. The factual correctness of summaries has the highest priority before practical applications.…

cs.CL202224 cited

Report from the NSF Future Directions Workshop on Automatic Evaluation of Dialog: Research Directions and Challenges

Shikib Mehri, Jinho Choi, Luis Fernando D'Haro +13

This is a report on the NSF Future Directions Workshop on Automatic Evaluation of Dialog. The workshop explored the current state of the art along with its limitations and suggeste…

cs.CL20228 cited

MDD-Eval: Self-Training on Augmented Data for Multi-Domain Dialogue Evaluation

Chen Zhang, Luis Fernando D'Haro, Thomas Friedrichs +1

Chatbots are designed to carry out human-like conversations across different domains, such as general chit-chat, knowledge exchange, and persona-grounded conversations. To measure…

cs.CL20211 cited

Investigating the Impact of Pre-trained Language Models on Dialog Evaluation

Chen Zhang, Luis Fernando D'Haro, Yiming Chen +2

Recently, there is a surge of interest in applying pre-trained language models (Pr-LM) in automatic open-domain dialog evaluation. Pr-LMs offer a promising direction for addressing…