3 papers
cs.CL2024
A Measure of the System Dependence of Automated Metrics
Pius von Däniken, Jan Deriu, Mark Cieliebak
Automated metrics for Machine Translation have made significant progress, with the goal of replacing expensive and time-consuming human evaluations. These metrics are typically ass…
cs.CL2022
On the Effectiveness of Automated Metrics for Text Generation Systems
Pius von Däniken, Jan Deriu, Don Tuggener +1
A major challenge in the field of Text Generation is evaluation because we lack a sound theory that can be leveraged to extract guidelines for evaluation campaigns. In this work, w…
cs.AI2022
Probing the Robustness of Trained Metrics for Conversational Dialogue Systems
Jan Deriu, Don Tuggener, Pius von Däniken +1
This paper introduces an adversarial method to stress-test trained metrics to evaluate conversational dialogue systems. The method leverages Reinforcement Learning to find response…