33 citations · 66 across the 11 of their papers we have counts for
14 papers · 1 filter
In the Blind: Building Pseudo-References for MT Evaluation
Diptesh Kanojia, Chi-kiu Lo, Archchana Sindhujan +3
The WMT26 General MT task evaluates systems on 10 language pairs that have no human references (neither translated from scratch nor post-edited from MT output by humans). We descri…
Dynamically Allocating Evaluation Effort for Model Ranking
Vilém Zouhar, Julia Kreutzer, Alon Lavie +4
While human evaluation is the gold standard in many NLP tasks, it suffers from prohibitive costs and poor scalability. When identifying top-performing models, typical evaluation pr…
Overview of Dialog System Evaluation Track: Dimensionality, Language, Culture and Safety at DSTC 12
John Mendonça, Lining Zhang, Rahul Mallidi +4
The rapid advancement of Large Language Models (LLMs) has intensified the need for robust dialogue system evaluation, yet comprehensive assessment remains challenging. Traditional…
MEDAL: A Framework for Benchmarking LLMs as Multilingual Open-Domain Dialogue Evaluators
John Mendonça, Alon Lavie, Isabel Trancoso
Evaluating the quality of open-domain chatbots has become increasingly reliant on LLMs acting as automatic judges. However, existing meta-evaluation benchmarks are static, outdated…
Soda-Eval: Open-Domain Dialogue Evaluation in the age of LLMs
John Mendonça, Isabel Trancoso, Alon Lavie
Although human evaluation remains the gold standard for open-domain dialogue evaluation, the growing popularity of automated evaluation using Large Language Models (LLMs) has also…
ECoh: Turn-level Coherence Evaluation for Multilingual Dialogues
John Mendonça, Isabel Trancoso, Alon Lavie
Despite being heralded as the new standard for dialogue evaluation, the closed-source nature of GPT-4 poses challenges for the community. Motivated by the need for lightweight, ope…