1 paper · 1 filter
Roy Xie, Junlin Wang, Ruomin Huang +5
The rapid scaling of large language models (LLMs) has raised concerns about the transparency and fair use of the data used in their pretraining. Detecting such content is challengi…