2 papers
cs.CL2026
Cultivar: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation Robustness
Pinzhen Chen, Koel Dutta Chowdhury, Xiaoya Xu +20
Multilingual translation benchmarks are typically sourced in English and translated into other languages, treating language pairs as the unit of evaluation---a design that is prone…
cs.CL2026
When Flores Bloomz Wrong: Cross-Direction Contamination in Machine Translation Evaluation
David Tan, Pinzhen Chen, Josef van Genabith +1
Large language models (LLMs) can be benchmark-contaminated, resulting in inflated scores that mask memorization as generalization, and in multilingual settings, this memorization c…