Showing cs.CLShow all
3 papers · 1 filter
cs.CL2024
Soda-Eval: Open-Domain Dialogue Evaluation in the age of LLMs
John Mendonça, Isabel Trancoso, Alon Lavie
Although human evaluation remains the gold standard for open-domain dialogue evaluation, the growing popularity of automated evaluation using Large Language Models (LLMs) has also…
cs.CL2024
ECoh: Turn-level Coherence Evaluation for Multilingual Dialogues
John Mendonça, Isabel Trancoso, Alon Lavie
Despite being heralded as the new standard for dialogue evaluation, the closed-source nature of GPT-4 poses challenges for the community. Motivated by the need for lightweight, ope…
cs.CL2024
On the Benchmarking of LLMs for Open-Domain Dialogue Evaluation
John Mendonça, Alon Lavie, Isabel Trancoso
Large Language Models (LLMs) have showcased remarkable capabilities in various Natural Language Processing tasks. For automatic open-domain dialogue evaluation in particular, LLMs…