4 papers
Robust Native Language Identification through Agentic Decomposition
Ahmet Yavuz Uluslu, Tannon Kew, Tilia Ellendorff +2
Large language models (LLMs) often achieve high performance in native language identification (NLI) benchmarks by leveraging superficial contextual clues such as names, locations,…
EMTeC: A Corpus of Eye Movements on Machine-Generated Texts
Lena Sophia Bolliger, Patrick Haller, Isabelle Caroline Rose Cretton +3
The Eye Movements on Machine-Generated Texts Corpus (EMTeC) is a naturalistic eye-movements-while-reading corpus of 107 native English speakers reading machine-generated texts. The…
Turning English-centric LLMs Into Polyglots: How Much Multilinguality Is Needed?
Tannon Kew, Florian Schottmann, Rico Sennrich
The vast majority of today's large language models (LLMs) are English-centric, having been pretrained predominantly on English text. Yet, in order to meet user expectations, models…
BLESS: Benchmarking Large Language Models on Sentence Simplification
Tannon Kew, Alison Chi, Laura Vásquez-Rodríguez +4
We present BLESS, a comprehensive performance benchmark of the most recent state-of-the-art large language models (LLMs) on the task of text simplification (TS). We examine how wel…