3 papers
cs.CL2026
CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data
Pedro Ortiz Suarez, Laurie Burchell, Catherine Arnett +94
Language identification (LID) is a fundamental step in curating multilingual corpora. However, LID models still perform poorly for many languages, especially on the noisy and heter…
cs.CY2026
A Human-Centric Framework for Data Attribution in Large Language Models
Amelie Wührl, Mattes Ruckdeschel, Kyle Lo +1
In the current Large Language Model (LLM) ecosystem, creators have little agency over how their data is used, and LLM users may find themselves unknowingly plagiarizing existing so…
cs.CL2024
PETapter: Leveraging PET-style classification heads for modular few-shot parameter-efficient fine-tuning
Jonas Rieger, Mattes Ruckdeschel, Gregor Wiedemann
Few-shot learning and parameter-efficient fine-tuning (PEFT) are crucial to overcome the challenges of data scarcity and ever growing language model sizes. This applies in particul…