activity
20202024
collaborators

5 papers

cs.AI2024

Favi-Score: A Measure for Favoritism in Automated Preference Ratings for Generative AI Evaluation

Pius von Däniken, Jan Deriu, Don Tuggener +1

Generative AI systems have become ubiquitous for all kinds of modalities, which makes the issue of the evaluation of such models more pressing. One popular approach is preference r…

cs.CL2023

Correction of Errors in Preference Ratings from Automated Metrics for Text Generation

Jan Deriu, Pius von Däniken, Don Tuggener +1

A major challenge in the field of Text Generation is evaluation: Human evaluations are cost-intensive, and automated metrics often display considerable disagreement with human judg…

cs.CL2022

On the Effectiveness of Automated Metrics for Text Generation Systems

Pius von Däniken, Jan Deriu, Don Tuggener +1

A major challenge in the field of Text Generation is evaluation because we lack a sound theory that can be leveraged to extract guidelines for evaluation campaigns. In this work, w…

cs.AI2022

Probing the Robustness of Trained Metrics for Conversational Dialogue Systems

Jan Deriu, Don Tuggener, Pius von Däniken +1

This paper introduces an adversarial method to stress-test trained metrics to evaluate conversational dialogue systems. The method leverages Reinforcement Learning to find response…

cs.AI2020

Spot The Bot: A Robust and Efficient Framework for the Evaluation of Conversational Dialogue Systems

Jan Deriu, Don Tuggener, Pius von Däniken +6

The lack of time-efficient and reliable evaluation methods hamper the development of conversational dialogue systems (chatbots). Evaluations requiring humans to converse with chatb…