3 citations · 3 across the 3 of their papers we have counts for
3 papers
cs.CL2024
Finding Replicable Human Evaluations via Stable Ranking Probability
Parker Riley, Daniel Deutsch, George Foster +3
Reliable human evaluation is critical to the development of successful natural language generation models, but achieving it is notoriously difficult. Stability is a crucial require…
cs.CL2024
To Diverge or Not to Diverge: A Morphosyntactic Perspective on Machine Translation vs Human Translation
Jiaming Luo, Colin Cherry, George Foster
We conduct a large-scale fine-grained comparative analysis of machine translations (MT) against human translations (HT) through the lens of morphosyntactic divergence. Across three…
cs.CL2023★ 3 cited
Ties Matter: Meta-Evaluating Modern Metrics with Pairwise Accuracy and Tie Calibration
Daniel Deutsch, George Foster, Markus Freitag
Kendall's tau is frequently used to meta-evaluate how well machine translation (MT) evaluation metrics score individual translations. Its focus on pairwise score comparisons is int…