1 citations · 1 across the 4 of their papers we have counts for
4 papers
On Leakage of Code Generation Evaluation Datasets
Alexandre Matton, Tom Sherborne, Dennis Aumiller +7
In this paper, we consider contamination by code generation test sets, in particular in their use in modern large language models. We discuss three possible sources of such contami…
BLESS: Benchmarking Large Language Models on Sentence Simplification
Tannon Kew, Alison Chi, Laura Vásquez-Rodríguez +4
We present BLESS, a comprehensive performance benchmark of the most recent state-of-the-art large language models (LLMs) on the task of text simplification (TS). We examine how wel…
Evaluating Factual Consistency of Texts with Semantic Role Labeling
Jing Fan, Dennis Aumiller, Michael Gertz
Automated evaluation of text generation systems has recently seen increasing attention, particularly checking whether generated text stays truthful to input sources. Existing metho…
On the State of German (Abstractive) Text Summarization
Dennis Aumiller, Jing Fan, Michael Gertz
With recent advancements in the area of Natural Language Processing, the focus is slowly shifting from a purely English-centric view towards more language-specific solutions, inclu…