activity
20172025
most citedContextualized Topic Coherence Metrics

11 citations · 24 across the 9 of their papers we have counts for

collaborators
Showing 2023Show all

5 papers · 1 filter

cs.CL2023

Data Similarity is Not Enough to Explain Language Model Performance

Gregory Yauney, Emily Reif, David Mimno

Large language models achieve high performance on many but not all downstream tasks. The interaction between pretraining data and task data is commonly assumed to determine this va…

cs.CL20231 cited

T5 meets Tybalt: Author Attribution in Early Modern English Drama Using Large Language Models

Rebecca M. M. Hicke, David Mimno

Large language models have shown breakthrough potential in many NLP domains. Here we consider their use for stylometry, specifically authorship identification in Early Modern Engli…

cs.CL20231 cited

Modeling Legal Reasoning: LM Annotation at the Edge of Human Agreement

Rosamond Thalken, Edward H. Stiglitz, David Mimno +1

Generative language models (LMs) are increasingly used for document class-prediction tasks and promise enormous improvements in cost and efficiency. Existing research often examine…

cs.CL202311 cited

Contextualized Topic Coherence Metrics

Hamed Rahimi, Jacob Louis Hoover, David Mimno +3

The recent explosion in work on neural topic modeling has been criticized for optimizing automated topic evaluation metrics at the expense of actual meaningful topic identification…

cs.CL2023

A Pretrainer's Guide to Training Data: Measuring the Effects of Data Age, Domain Coverage, Quality, & Toxicity

Shayne Longpre, Gregory Yauney, Emily Reif +8

Pretraining is the preliminary and fundamental step in developing capable language models (LM). Despite this, pretraining data design is critically under-documented and often guide…