5 papers · 1 filter
LLMs as Span Annotators: A Comparative Study of LLMs and Humans
Zdeněk Kasner, Vilém Zouhar, Patrícia Schmidtová +7
Span annotation - annotating specific text features at the span level - can be used to evaluate texts where single-score metrics fail to provide actionable feedback. Until recently…
Automatic Metrics in Natural Language Generation: A Survey of Current Evaluation Practices
Patrícia Schmidtová, Saad Mahamood, Simone Balloccu +6
Automatic metrics are extensively used to evaluate natural language processing systems. However, there has been increasing focus on how they are used and reported by practitioners…
factgenie: A Framework for Span-based Evaluation of Generated Texts
Zdeněk Kasner, Ondřej Plátek, Patrícia Schmidtová +2
We present factgenie: a framework for annotating and visualizing word spans in textual model outputs. Annotations can capture various span-based phenomena such as semantic inaccura…
With a Little Help from the Authors: Reproducing Human Evaluation of an MT Error Detector
Ondřej Plátek, Mateusz Lango, Ondřej Dušek
This work presents our efforts to reproduce the results of the human evaluation experiment presented in the paper of Vamvas and Sennrich (2022), which evaluated an automatic system…
Three Ways of Using Large Language Models to Evaluate Chat
Ondřej Plátek, Vojtěch Hudeček, Patricia Schmidtová +2
This paper describes the systems submitted by team6 for ChatEval, the DSTC 11 Track 4 competition. We present three different approaches to predicting turn-level qualities of chatb…