Showing cs.AIShow all
2 papers · 1 filter
cs.AI2024
Favi-Score: A Measure for Favoritism in Automated Preference Ratings for Generative AI Evaluation
Pius von Däniken, Jan Deriu, Don Tuggener +1
Generative AI systems have become ubiquitous for all kinds of modalities, which makes the issue of the evaluation of such models more pressing. One popular approach is preference r…
cs.AI2022
Probing the Robustness of Trained Metrics for Conversational Dialogue Systems
Jan Deriu, Don Tuggener, Pius von Däniken +1
This paper introduces an adversarial method to stress-test trained metrics to evaluate conversational dialogue systems. The method leverages Reinforcement Learning to find response…