activity
20152025
most citedTowards Using Machine Translation Techniques to Induce Multilingual Lexica of Discourse Markers

6 citations · 20 across the 13 of their papers we have counts for

collaborators
Showing cs.CLShow all

8 papers · 1 filter

cs.CL2025

MEDAL: A Framework for Benchmarking LLMs as Multilingual Open-Domain Dialogue Evaluators

John Mendonça, Alon Lavie, Isabel Trancoso

Evaluating the quality of open-domain chatbots has become increasingly reliant on LLMs acting as automatic judges. However, existing meta-evaluation benchmarks are static, outdated…

cs.CL2024

Soda-Eval: Open-Domain Dialogue Evaluation in the age of LLMs

John Mendonça, Isabel Trancoso, Alon Lavie

Although human evaluation remains the gold standard for open-domain dialogue evaluation, the growing popularity of automated evaluation using Large Language Models (LLMs) has also…

cs.CL2024

ECoh: Turn-level Coherence Evaluation for Multilingual Dialogues

John Mendonça, Isabel Trancoso, Alon Lavie

Despite being heralded as the new standard for dialogue evaluation, the closed-source nature of GPT-4 poses challenges for the community. Motivated by the need for lightweight, ope…

cs.CL2024★ 1 cited

On the Benchmarking of LLMs for Open-Domain Dialogue Evaluation

John Mendonça, Alon Lavie, Isabel Trancoso

Large Language Models (LLMs) have showcased remarkable capabilities in various Natural Language Processing tasks. For automatic open-domain dialogue evaluation in particular, LLMs…

cs.CL2023

Dialogue Quality and Emotion Annotations for Customer Support Conversations

John Mendonça, Patrícia Pereira, Miguel Menezes +6

Task-oriented conversational datasets often lack topic variability and linguistic diversity. However, with the advent of Large Language Models (LLMs) pretrained on extensive, multi…

cs.CL2023★ 6 cited

Simple LLM Prompting is State-of-the-Art for Robust and Multilingual Dialogue Evaluation

John Mendonça, Patrícia Pereira, Helena Moniz +3

Despite significant research effort in the development of automatic dialogue evaluation metrics, little thought is given to evaluating dialogues other than in English. At the same…