1 paper
Xiao Yang, Erik Edward Aldape, Beren Millidge
Large language model training corpora grow through successive, often redundant releases, so each release must be deduplicated against both itself and the accumulated history. At tr…