most citedLeak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs

19 citations · 20 across the 5 of their papers we have counts for

collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2024

Automatic Metrics in Natural Language Generation: A Survey of Current Evaluation Practices

Patrícia Schmidtová, Saad Mahamood, Simone Balloccu +6

Automatic metrics are extensively used to evaluate natural language processing systems. However, there has been increasing focus on how they are used and reported by practitioners…

cs.CL2024

factgenie: A Framework for Span-based Evaluation of Generated Texts

Zdeněk Kasner, Ondřej Plátek, Patrícia Schmidtová +2

We present factgenie: a framework for annotating and visualizing word spans in textual model outputs. Annotations can capture various span-based phenomena such as semantic inaccura…

cs.CL202419 cited

Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs

Simone Balloccu, Patrícia Schmidtová, Mateusz Lango +1

Natural Language Processing (NLP) research is increasingly focusing on the use of Large Language Models (LLMs), with some of the most popular ones being either fully or partially c…

cs.CL20241 cited

Ask the experts: sourcing high-quality datasets for nutritional counselling through Human-AI collaboration

Simone Balloccu, Ehud Reiter, Vivek Kumar +2

Large Language Models (LLMs), with their flexible generation abilities, can be powerful data sources in domains with few or no available corpora. However, problems like hallucinati…

cs.CL2022

Comparing informativeness of an NLG chatbot vs graphical app in diet-information domain

Simone Balloccu, Ehud Reiter

Visual representation of data like charts and tables can be challenging to understand for readers. Previous work showed that combining visualisations with text can improve the comm…