Showing cs.AIShow all
2 papers · 1 filter
cs.AI2022
Probing the Robustness of Trained Metrics for Conversational Dialogue Systems
Jan Deriu, Don Tuggener, Pius von Däniken +1
This paper introduces an adversarial method to stress-test trained metrics to evaluate conversational dialogue systems. The method leverages Reinforcement Learning to find response…
cs.AI2020
Spot The Bot: A Robust and Efficient Framework for the Evaluation of Conversational Dialogue Systems
Jan Deriu, Don Tuggener, Pius von Däniken +6
The lack of time-efficient and reliable evaluation methods hamper the development of conversational dialogue systems (chatbots). Evaluations requiring humans to converse with chatb…