4 citations · 6 across the 5 of their papers we have counts for
5 papers
Automatic Metrics in Natural Language Generation: A Survey of Current Evaluation Practices
Patrícia Schmidtová, Saad Mahamood, Simone Balloccu +6
Automatic metrics are extensively used to evaluate natural language processing systems. However, there has been increasing focus on how they are used and reported by practitioners…
factgenie: A Framework for Span-based Evaluation of Generated Texts
Zdeněk Kasner, Ondřej Plátek, Patrícia Schmidtová +2
We present factgenie: a framework for annotating and visualizing word spans in textual model outputs. Annotations can capture various span-based phenomena such as semantic inaccura…
With a Little Help from the Authors: Reproducing Human Evaluation of an MT Error Detector
Ondřej Plátek, Mateusz Lango, Ondřej Dušek
This work presents our efforts to reproduce the results of the human evaluation experiment presented in the paper of Vamvas and Sennrich (2022), which evaluated an automatic system…
Three Ways of Using Large Language Models to Evaluate Chat
Ondřej Plátek, Vojtěch Hudeček, Patricia Schmidtová +2
This paper describes the systems submitted by team6 for ChatEval, the DSTC 11 Track 4 competition. We present three different approaches to predicting turn-level qualities of chatb…
Recurrent Neural Networks for Dialogue State Tracking
Ondřej Plátek, Petr Bělohlávek, Vojtěch Hudeček +1
This paper discusses models for dialogue state tracking using recurrent neural networks (RNN). We present experiments on the standard dialogue state tracking (DST) dataset, DSTC2.…