10 citations · 19 across the 6 of their papers we have counts for
6 papers
FineD-Eval: Fine-grained Automatic Dialogue-Level Evaluation
Chen Zhang, Luis Fernando D'Haro, Qiquan Zhang +2
Recent model-based reference-free metrics for open-domain dialogue evaluation exhibit promising correlations with human judgment. However, they either perform turn-level evaluation…
A Focused Study on Sequence Length for Dialogue Summarization
Bin Wang, Chen Zhang, Chengwei Wei +1
Output length is critical to dialogue summarization systems. The dialogue summary length is determined by multiple factors, including dialogue complexity, summary objective, and pe…
Analyzing and Evaluating Faithfulness in Dialogue Summarization
Bin Wang, Chen Zhang, Yan Zhang +2
Dialogue summarization is abstractive in nature, making it suffer from factual errors. The factual correctness of summaries has the highest priority before practical applications.…
Just Rank: Rethinking Evaluation with Word and Sentence Similarities
Bin Wang, C. -C. Jay Kuo, Haizhou Li
Word and sentence embeddings are useful feature representations in natural language processing. However, intrinsic evaluation for embeddings lags far behind, and there has been no…
MDD-Eval: Self-Training on Augmented Data for Multi-Domain Dialogue Evaluation
Chen Zhang, Luis Fernando D'Haro, Thomas Friedrichs +1
Chatbots are designed to carry out human-like conversations across different domains, such as general chit-chat, knowledge exchange, and persona-grounded conversations. To measure…
Investigating the Impact of Pre-trained Language Models on Dialog Evaluation
Chen Zhang, Luis Fernando D'Haro, Yiming Chen +2
Recently, there is a surge of interest in applying pre-trained language models (Pr-LM) in automatic open-domain dialog evaluation. Pr-LMs offer a promising direction for addressing…