2 papers
cs.CL2026
Dynamically Allocating Evaluation Effort for Model Ranking
Vilém Zouhar, Vilém Zouhar, Julia Kreutzer +6
While human evaluation is the gold standard in many NLP tasks, it suffers from prohibitive costs and poor scalability. When identifying top-performing models, typical evaluation pr…
cs.CL2026
PEAR: Pairwise Evaluation for Automatic Relative Scoring in Machine Translation
Lorenzo Proietti, Roman Grundkiewicz, Matt Post
We present PEAR (Pairwise Evaluation for Automatic Relative Scoring), a supervised quality estimation (QE) metric family that reframes reference-free machine translation (MT) evalu…