collaborators

5 papers

cs.CL2026

How Good is Your Wikipedia? Auditing Data Quality for Low-resource and Multilingual NLP

Kushal Tatariya, Artur Kulmizev, Wessel Poelman +6

Wikipedia's perceived high quality and broad language coverage have established it as a fundamental resource in NLP. However, in recent years, such assumptions of high quality have…

cs.CL2025

SMOL: Professionally translated parallel data for 115 under-represented languages

Isaac Caswell, Elizabeth Nielsen, Jiaming Luo +23

We open-source SMOL (Set of Maximal Overall Leverage), a suite of training data to unlock machine translation for low-resource languages. SMOL has been translated into 124 (and gro…

cs.CL2025

Gemma 3 Technical Report

Gemma Team, Aishwarya Kamath, Johan Ferret +209

We introduce Gemma 3, a multimodal addition to the Gemma family of lightweight open models, ranging in scale from 1 to 27 billion parameters. This version introduces vision underst…

cs.CL2025

Leveraging Domain Knowledge at Inference Time for LLM Translation: Retrieval versus Generation

Bryan Li, Jiaming Luo, Eleftheria Briakou +1

While large language models (LLMs) have been increasingly adopted for machine translation (MT), their performance for specialist domains such as medicine and law remains an open ch…

cs.CL2025

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation

Muhammed Yusuf Kocyigit, Eleftheria Briakou, Daniel Deutsch +3

Data contamination -- the accidental consumption of evaluation examples within the pre-training data -- can undermine the validity of evaluation benchmarks. In this paper, we prese…