papers
Publications (2)
cs.LG2024
Perplexed by Perplexity: Perplexity-Based Data Pruning With Small Reference Models
Zachary Ankner, Cody Blakeney, Kartik Sreenivasan +3
In this work, we investigate whether small language models can determine high-quality subsets of large-scale text datasets that improve the performance of larger language models. W…
cs.CL2023
When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale
Max Marion, Ahmet Ãstün, Luiza Pozzobon +3
Large volumes of text data have contributed significantly to the development of large language models (LLMs) in recent years. This data is typically acquired by scraping the intern…