65 citations · 181 across the 17 of their papers we have counts for
Showing 2024 · cs.CLShow all
2 papers · 2 filters
cs.CL2024★ 3 cited
Improving Pretraining Data Using Perplexity Correlations
Tristan Thrush, Christopher Potts, Tatsunori Hashimoto
Quality pretraining data is often seen as the key to high-performance language models. However, progress in understanding pretraining data has been slow due to the costly pretraini…
cs.CL2024
I am a Strange Dataset: Metalinguistic Tests for Language Models
Tristan Thrush, Jared Moore, Miguel Monares +2
Statements involving metalinguistic self-reference ("This paper has six sections.") are prevalent in many domains. Can current large language models (LLMs) handle such language? In…