From the 1 of 15 linked papers with an AI index.
15 papers
There is No Theoretical Curse of Multilinguality For Embedding Space Structure
Niyati Bafna, Neha Verma, Vilém Zouhar +2
A central goal of multilingual NLP is to achieve high monolingual performance per language and cross-lingual alignment for large-scale language coverage with a multilingual model.…
Cultivar: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation Robustness
Pinzhen Chen, Koel Dutta Chowdhury, Xiaoya Xu +20
Multilingual translation benchmarks are typically sourced in English and translated into other languages, treating language pairs as the unit of evaluation---a design that is prone…
Dynamically Allocating Evaluation Effort for Model Ranking
Vilém Zouhar, Vilém Zouhar, Julia Kreutzer +6
While human evaluation is the gold standard in many NLP tasks, it suffers from prohibitive costs and poor scalability. When identifying top-performing models, typical evaluation pr…
Contrastive ESA: Human Evaluation of Multiple Translations at Once
Vilém Zouhar, Roman Grundkiewicz, Sara Rajaee +6
The paper proposes Contrastive Error Span Annotation (cESA), a human evaluation protocol that shows multiple translations of the same source together, lets annotators mark error sp…
Searching the Internet for Challenging Benchmarks at Scale
Wenda Xu, Vilém Zouhar, Parker Riley +3
Many static benchmarks are beginning to saturate: as models rapidly improve, they achieve near-perfect scores on fixed test sets, leaving little headroom to expose genuine model we…
When LLMs Benchmark Themselves: Deconstructing Self-Bias in Automated Evaluation
Wenda Xu, Sweta Agrawal, Vilém Zouhar +2
As LLMs rapidly saturate existing benchmarks, automated benchmark creation using LLMs (LLM-as-a-benchmark) -- where a model generates test inputs (LLM-as-a-testset) and evaluates o…