activity
20212025
most citedThree Ways of Using Large Language Models to Evaluate Chat

2 citations · 2 across the 4 of their papers we have counts for

collaborators

6 papers

cs.CL2025

How Important is `Perfect' English for Machine Translation Prompts?

Patrícia Schmidtová, Niyati Bafna, Seth Aycock +4

Large language models (LLMs) have achieved top results in recent machine translation evaluations, but they are also known to be sensitive to errors and perturbations in their promp…

cs.CL2025

LLMs as Span Annotators: A Comparative Study of LLMs and Humans

Zdeněk Kasner, Vilém Zouhar, Patrícia Schmidtová +7

Span annotation - annotating specific text features at the span level - can be used to evaluate texts where single-score metrics fail to provide actionable feedback. Until recently…

cs.CL2024

Automatic Metrics in Natural Language Generation: A Survey of Current Evaluation Practices

Patrícia Schmidtová, Saad Mahamood, Simone Balloccu +6

Automatic metrics are extensively used to evaluate natural language processing systems. However, there has been increasing focus on how they are used and reported by practitioners…

cs.CL2024

factgenie: A Framework for Span-based Evaluation of Generated Texts

Zdeněk Kasner, Ondřej Plátek, Patrícia Schmidtová +2

We present factgenie: a framework for annotating and visualizing word spans in textual model outputs. Annotations can capture various span-based phenomena such as semantic inaccura…

cs.CL20232 cited

Three Ways of Using Large Language Models to Evaluate Chat

Ondřej Plátek, Vojtěch Hudeček, Patricia Schmidtová +2

This paper describes the systems submitted by team6 for ChatEval, the DSTC 11 Track 4 competition. We present three different approaches to predicting turn-level qualities of chatb…

cs.CL2021

THEaiTRE 1.0: Interactive generation of theatre play scripts

Rudolf Rosa, Tomáš Musil, Ondřej Dušek +13

We present the first version of a system for interactive generation of theatre play scripts. The system is based on a vanilla GPT-2 model with several adjustments, targeting specif…