activity
20212025
most citedWhen Scaling Meets LLM Finetuning: The Effect of Data, Model and Finetuning Method

28 citations · 57 across the 7 of their papers we have counts for

collaborators
Showing cs.CLShow all

7 papers · 1 filter

cs.CL20262 cited

TranslateGemma Technical Report

Mara Finkelstein, Isaac Caswell, Tobias Domhan +18

We present TranslateGemma, a suite of open machine translation models based on the Gemma 3 foundation models. To enhance the inherent multilingual capabilities of Gemma 3 for the t…

cs.CL2024

Don't Throw Away Data: Better Sequence Knowledge Distillation

Jun Wang, Eleftheria Briakou, Hamid Dadkhahi +3

A critical component in knowledge distillation is the means of coupling the teacher and student. The predominant sequence knowledge distillation method involves supervised learning…

cs.CL202428 cited

When Scaling Meets LLM Finetuning: The Effect of Data, Model and Finetuning Method

Biao Zhang, Zhongtao Liu, Colin Cherry +1

While large language models (LLMs) often adopt finetuning to unlock their capabilities for downstream applications, our understanding on the inductive biases (especially the scalin…

cs.CL2024

To Diverge or Not to Diverge: A Morphosyntactic Perspective on Machine Translation vs Human Translation

Jiaming Luo, Colin Cherry, George Foster

We conduct a large-scale fine-grained comparative analysis of machine translations (MT) against human translations (HT) through the lens of morphosyntactic divergence. Across three…

cs.CL20234 cited

Searching for Needles in a Haystack: On the Role of Incidental Bilingualism in PaLM's Translation Capability

Eleftheria Briakou, Colin Cherry, George Foster

Large, multilingual language models exhibit surprisingly good zero- or few-shot machine translation capabilities, despite having never seen the intentionally-included translation e…

cs.CL202324 cited

The unreasonable effectiveness of few-shot learning for machine translation

Xavier Garcia, Yamini Bansal, Colin Cherry +5

We demonstrate the potential of few-shot translation systems, trained with unpaired language data, for both high and low-resource language pairs. We show that with only 5 examples…