2 papers
cs.CL2026
FLUX: Data Worth Training On
Gowtham, Sai Rupesh, Sanjay Kumar +2
Modern large language model training is no longer limited by data availability, but by the inability of existing preprocessing pipelines to simultaneously achieve massive scale and…
cs.CL2025
Blu-WERP (Web Extraction and Refinement Pipeline): A Scalable Pipeline for Preprocessing Large Language Model Datasets
Gowtham, Sai Rupesh, Sanjay Kumar +2
High-quality training data is fundamental to large language model (LLM) performance, yet existing preprocessing pipelines often struggle to effectively remove noise and unstructure…