Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Does Training on Synthetic Data Make Models Less Robust?
Lingze Zhang, Ellie Pavlick
An increasingly common practice is to train large language models (LLMs) using synthetic data. Often this synthetic data is produced by the same or similar LLMs as those it is bein…
cs.CL2025
Doc-Guided Sent2Sent++: A Sent2Sent++ Agent with Doc-Guided memory for Document-level Machine Translation
Jiaxin Guo, Yuanchang Luo, Daimeng Wei +8
The field of artificial intelligence has witnessed significant advancements in natural language processing, largely attributed to the capabilities of Large Language Models (LLMs).…