collaborators

6 papers

cs.CL2026

MEDAL: A Framework for Benchmarking LLMs as Multilingual Open-Domain Dialogue Evaluators

John Mendonça, Alon Lavie, Isabel Trancoso

Evaluating the quality of open-domain chatbots has become increasingly reliant on LLMs acting as automatic judges. However, existing meta-evaluation benchmarks are static, outdated…

cs.CL2024

Soda-Eval: Open-Domain Dialogue Evaluation in the age of LLMs

John Mendonça, Isabel Trancoso, Alon Lavie

Although human evaluation remains the gold standard for open-domain dialogue evaluation, the growing popularity of automated evaluation using Large Language Models (LLMs) has also…

eess.AS2024

Privacy-oriented manipulation of speaker representations

Francisco Teixeira, Alberto Abad, Bhiksha Raj +1

Speaker embeddings are ubiquitous, with applications ranging from speaker recognition and diarization to speech synthesis and voice anonymisation. The amount of information held by…

cs.CL2024

ECoh: Turn-level Coherence Evaluation for Multilingual Dialogues

John Mendonça, Isabel Trancoso, Alon Lavie

Despite being heralded as the new standard for dialogue evaluation, the closed-source nature of GPT-4 poses challenges for the community. Motivated by the need for lightweight, ope…

cs.CL2024

On the Benchmarking of LLMs for Open-Domain Dialogue Evaluation

John Mendonça, Alon Lavie, Isabel Trancoso

Large Language Models (LLMs) have showcased remarkable capabilities in various Natural Language Processing tasks. For automatic open-domain dialogue evaluation in particular, LLMs…

cs.LG2024

Improving Membership Inference in ASR Model Auditing with Perturbed Loss Features

Francisco Teixeira, Karla Pizzi, Raphael Olivier +3

Membership Inference (MI) poses a substantial privacy threat to the training data of Automatic Speech Recognition (ASR) systems, while also offering an opportunity to audit these m…