6 papers · 1 filter
Tiny Aya: Bridging Scale and Multilingual Depth
Alejandro R. Salamanca, Diana Abagyan, Daniel D'souza +23
Tiny Aya redefines what a small multilingual language model can achieve. Trained on 70 languages and refined through region-aware posttraining, it delivers state-of-the-art in tran…
AI-Assisted Human Evaluation of Machine Translation
Vilém Zouhar, Tom Kocmi, Mrinmaya Sachan
Annually, research teams spend large amounts of money to evaluate the quality of machine translation systems (WMT, inter alia). This is expensive because it requires a lot of exper…
Error Span Annotation: A Balanced Approach for Human Evaluation of Machine Translation
Tom Kocmi, Vilém Zouhar, Eleftherios Avramidis +5
High-quality Machine Translation (MT) evaluation relies heavily on human judgments. Comprehensive error classification methods, such as Multidimensional Quality Metrics (MQM), are…
Preliminary WMT24 Ranking of General MT Systems and LLMs
Tom Kocmi, Eleftherios Avramidis, Rachel Bawden +18
This is the preliminary ranking of WMT24 General MT systems based on automatic metrics. The official ranking will be a human evaluation, which is superior to the automatic ranking…
Navigating the Metrics Maze: Reconciling Score Magnitudes and Accuracies
Tom Kocmi, Vilém Zouhar, Christian Federmann +1
Ten years ago a single metric, BLEU, governed progress in machine translation research. For better or worse, there is no such consensus today, and consequently it is difficult for…
Not All Metrics Are Guilty: Improving NLG Evaluation by Diversifying References
Tianyi Tang, Hongyuan Lu, Yuchen Eleanor Jiang +5
Most research about natural language generation (NLG) relies on evaluation benchmarks with limited references for a sample, which may result in poor correlations with human judgeme…