Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
MEDAL: A Framework for Benchmarking LLMs as Multilingual Open-Domain Dialogue Evaluators
John Mendonça, Alon Lavie, Isabel Trancoso
Evaluating the quality of open-domain chatbots has become increasingly reliant on LLMs acting as automatic judges. However, existing meta-evaluation benchmarks are static, outdated…
cs.CL2025
Overview of Dialog System Evaluation Track: Dimensionality, Language, Culture and Safety at DSTC 12
John Mendonça, Lining Zhang, Rahul Mallidi +4
The rapid advancement of Large Language Models (LLMs) has intensified the need for robust dialogue system evaluation, yet comprehensive assessment remains challenging. Traditional…
cs.CL2025
CAMÃES: A Comprehensive Automatic Speech Recognition Benchmark for European Portuguese
Carlos Carvalho, Francisco Teixeira, Catarina Botelho +9
Existing resources for Automatic Speech Recognition in Portuguese are mostly focused on Brazilian Portuguese, leaving European Portuguese (EP) and other varieties under-explored. T…