9 papers · 1 filter
From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set
Mara Finkelstein, Dan Deutsch, Parker Riley +3
As LLMs continue to become more powerful and versatile, human evaluation has quickly become intractable at scale and reliance on automatic metrics has become the norm. Recently, it…
Introducing the NewsPaLM MBR and QE Dataset: LLM-Generated High-Quality Parallel Data Outperforms Traditional Web-Crawled Data
Mara Finkelstein, David Vilar, Markus Freitag
Recent research in neural machine translation (NMT) has shown that training on high-quality machine-generated data can outperform training on human-generated data. This work accomp…
Mitigating Metric Bias in Minimum Bayes Risk Decoding
Geza Kovacs, Daniel Deutsch, Markus Freitag
While Minimum Bayes Risk (MBR) decoding using metrics such as COMET or MetricX has outperformed traditional decoding methods such as greedy or beam search, it introduces a challeng…
LLMRefine: Pinpointing and Refining Large Language Models via Fine-Grained Actionable Feedback
Wenda Xu, Daniel Deutsch, Mara Finkelstein +6
Recent large language models (LLM) are leveraging human feedback to improve their generation quality. However, human feedback is costly to obtain, especially during inference. In t…
Learning from others' mistakes: Finetuning machine translation models with span-level error annotations
Lily H. Zhang, Hamid Dadkhahi, Mara Finkelstein +3
Despite growing interest in incorporating feedback to improve language models, most efforts focus only on sequence-level annotations. In this work, we explore the potential of util…
Beyond Human-Only: Evaluating Human-Machine Collaboration for Collecting High-Quality Translation Data
Zhongtao Liu, Parker Riley, Daniel Deutsch +4
Collecting high-quality translations is crucial for the development and evaluation of machine translation systems. However, traditional human-only approaches are costly and slow. T…