1 citations · 1 across the 12 of their papers we have counts for
1 paper · 2 filters
Nan He, Weichen Xiong, Hanwen Liu +6
The effectiveness of large language models (LLMs) is often hindered by duplicated data in their extensive pre-training datasets. Current approaches primarily focus on detecting and…