10 papers
Rewrite to Translate, Translate to Reward: Reinforcement Learning for Source Rewriting in Machine Translation
Boxuan Lyu, Haiyue Song, Zhi Qu +3
Prior work has explored prompting large language models (LLMs) to rewrite source text before translation, with the goal of improving machine translation (MT) quality. However, we f…
ExCAM: Explainable Cultural Awareness Metrics
Christoph Leiter, Haiyue Song, Hour Kaing +4
Evaluating the cultural awareness of large language models is crucial to ensure the fairness of generated text and the generalizability of applications across the world. Recent ben…
OptiMer: Optimal Distribution Vector Merging Is Better than Data Mixing for Continual Pre-Training
Haiyue Song, Masao Utiyama
Continual pre-training is widely used to adapt LLMs to target languages and domains, yet the mixture ratio of training data remains a sensitive hyperparameter that is expensive to…
Is Human Annotation Necessary? Iterative MBR Distillation for Error Span Detection in Machine Translation
Boxuan Lyu, Haiyue Song, Zhi Qu
Error Span Detection (ESD) is a crucial subtask in Machine Translation (MT) evaluation, aiming to identify the location and severity of translation errors. While fine-tuning models…
PrahokBART: A Pre-trained Sequence-to-Sequence Model for Khmer Natural Language Generation
Hour Kaing, Raj Dabre, Haiyue Song +3
This work introduces {\it PrahokBART}, a compact pre-trained sequence-to-sequence model trained from scratch for Khmer using carefully curated Khmer and English corpora. We focus o…
Structured Document Translation via Format Reinforcement Learning
Haiyue Song, Johannes Eschbach-Dymanus, Hour Kaing +4
Recent works on structured text translation remain limited to the sentence level, as they struggle to effectively handle the complex document-level XML or HTML structures. To addre…