3 papers
cs.CL2025
Token-level Ensembling of Models with Different Vocabularies
Rachel Wicks, Kartik Ravisankar, Xinchen Yang +2
Model ensembling is a technique to combine the predicted distributions of two or more models, often leading to improved robustness and performance. For ensembling in text generatio…
cs.CL2024
Recovering document annotations for sentence-level bitext
Rachel Wicks, Matt Post, Philipp Koehn
Data availability limits the scope of any given task. In machine translation, historical models were incapable of handling longer contexts, so the lack of document-level datasets w…
cs.CL2023
Identifying Context-Dependent Translations for Evaluation Set Production
Rachel Wicks, Matt Post
A major impediment to the transition to context-aware machine translation is the absence of good evaluation metrics and test sets. Sentences that require context to be translated c…