6 citations · 8 across the 4 of their papers we have counts for
4 papers
Simple LLM Prompting is State-of-the-Art for Robust and Multilingual Dialogue Evaluation
John Mendonça, Patrícia Pereira, Helena Moniz +3
Despite significant research effort in the development of automatic dialogue evaluation metrics, little thought is given to evaluating dialogues other than in English. At the same…
Towards Multilingual Automatic Dialogue Evaluation
John Mendonça, Alon Lavie, Isabel Trancoso
The main limiting factor in the development of robust multilingual dialogue evaluation metrics is the lack of multilingual data and the limited availability of open sourced multili…
The Inside Story: Towards Better Understanding of Machine Translation Neural Evaluation Metrics
Ricardo Rei, Nuno M. Guerreiro, Marcos Treviso +3
Neural metrics for machine translation evaluation, such as COMET, exhibit significant improvements in their correlation with human judgments, as compared to traditional metrics bas…
Appropriateness is all you need!
Hendrik Kempt, Alon Lavie, Saskia K. Nagel
The strive to make AI applications "safe" has led to the development of safety-measures as the main or even sole normative requirement of their permissible use. Similar can be atte…