2 citations · 3 across the 3 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2024★ 1 cited
SoftDedup: an Efficient Data Reweighting Method for Speeding Up Language Model Pre-training
Nan He, Weichen Xiong, Hanwen Liu +6
The effectiveness of large language models (LLMs) is often hindered by duplicated data in their extensive pre-training datasets. Current approaches primarily focus on detecting and…
cs.CL2023
Zero-shot Cross-lingual Transfer without Parallel Corpus
Yuyang Zhang, Xiaofeng Han, Baojun Wang
Recently, although pre-trained language models have achieved great success on multilingual NLP (Natural Language Processing) tasks, the lack of training data on many tasks in low-r…