3 papers
cs.DC2026
FOLD: Fuzzy Online Deduplication for Very Large Evolving Datasets via Approximate Nearest Neighbor Search
Nelson Bore, Pritish Mishra, Constantin Adam +2
Fuzzy deduplication is key to constructing large language model training corpora. However, classic Locality-Sensitive Hashing (LSH) pipelines scale poorly as corpora grow and are i…
cs.CL2025
GneissWeb: Preparing High Quality Data for LLMs at Scale
Hajar Emami Gohari, Swanand Ravindra Kadhe, Syed Yousaf Shah +29
Data quantity and quality play a vital role in determining the performance of Large Language Models (LLMs). High-quality data, in particular, can significantly boost the LLM's abil…
cs.AI2024
Data-Prep-Kit: getting your data ready for LLM application development
David Wood, Boris Lublinsky, Alexy Roytman +21
Data preparation is the first and a very important step towards any Large Language Model (LLM) development. This paper introduces an easy-to-use, extensible, and scale-flexible ope…