Showing cs.CLShow all
2 papers · 1 filter
cs.CL2024
The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale
Guilherme Penedo, Hynek KydlÃÄek, Loubna Ben allal +5
The performance of a large language model (LLM) depends heavily on the quality and size of its pretraining dataset. However, the pretraining datasets for state-of-the-art open LLMs…
cs.CL2024
Constructing the CORD-19 Vaccine Dataset
Manisha Singh, Divy Sharma, Alonso Ma +2
We introduce new dataset 'CORD-19-Vaccination' to cater to scientists specifically looking into COVID-19 vaccine-related research. This dataset is extracted from CORD-19 dataset [W…