11 citations · 24 across the 9 of their papers we have counts for
5 papers · 1 filter
Data Similarity is Not Enough to Explain Language Model Performance
Gregory Yauney, Emily Reif, David Mimno
Large language models achieve high performance on many but not all downstream tasks. The interaction between pretraining data and task data is commonly assumed to determine this va…
T5 meets Tybalt: Author Attribution in Early Modern English Drama Using Large Language Models
Rebecca M. M. Hicke, David Mimno
Large language models have shown breakthrough potential in many NLP domains. Here we consider their use for stylometry, specifically authorship identification in Early Modern Engli…
Modeling Legal Reasoning: LM Annotation at the Edge of Human Agreement
Rosamond Thalken, Edward H. Stiglitz, David Mimno +1
Generative language models (LMs) are increasingly used for document class-prediction tasks and promise enormous improvements in cost and efficiency. Existing research often examine…
Contextualized Topic Coherence Metrics
Hamed Rahimi, Jacob Louis Hoover, David Mimno +3
The recent explosion in work on neural topic modeling has been criticized for optimizing automated topic evaluation metrics at the expense of actual meaningful topic identification…
A Pretrainer's Guide to Training Data: Measuring the Effects of Data Age, Domain Coverage, Quality, & Toxicity
Shayne Longpre, Gregory Yauney, Emily Reif +8
Pretraining is the preliminary and fundamental step in developing capable language models (LM). Despite this, pretraining data design is critically under-documented and often guide…