papers

Publications (11)

cs.CV2019

Computational Histological Staining and Destaining of Prostate Core Biopsy RGB Images with Generative Adversarial Neural Networks

Aman Rana, Gregory Yauney, Alarice Lowe +1

Histopathology tissue samples are widely available in two states: paraffin-embedded unstained and non-paraffin-embedded stained whole slide RGB images (WSRI). Hematoxylin and eosin…

cs.CL2026

BenchBrowser: Retrieving Evidence for Evaluating Benchmark Validity

Harshita Diddee, Gregory Yauney, Swabha Swayamdipta +1

Do language model benchmarks actually measure what practitioners intend them to ? High-level metadata is too coarse to convey the granular reality of benchmarks: a "poetry" benchma…

cs.CL2025

Sample, Align, Synthesize: Graph-Based Response Synthesis with ConGrs

Sayan Ghosh, Shahzaib Saqib Warraich, Dhruv Tarsadiya +2

Language models can be sampled multiple times to access the distribution underlying their responses, but existing methods cannot efficiently synthesize rich epistemic signals acros…

cs.CL2021

Comparing Text Representations: A Theory-Driven Approach

Gregory Yauney, David Mimno

Much of the progress in contemporary NLP has come from learning representations, such as masked language model (MLM) contextual embeddings, that turn challenging problems into simp…

cs.CL2023

Data Similarity is Not Enough to Explain Language Model Performance

Gregory Yauney, Emily Reif, David Mimno

Large language models achieve high performance on many but not all downstream tasks. The interaction between pretraining data and task data is commonly assumed to determine this va…

cs.CL2024

Stronger Random Baselines for In-Context Learning

Gregory Yauney, David Mimno

Evaluating the in-context learning classification performance of language models poses challenges due to small dataset sizes, extensive prompt-selection using the validation set, a…

cs.LG2018

Automated Process Incorporating Machine Learning Segmentation and Correlation of Oral Diseases with Systemic Health

Gregory Yauney, Aman Rana, Lawrence C. Wong +3

Imaging fluorescent disease biomarkers in tissues and skin is a non-invasive method to screen for health conditions. We report an automated process that combines intraoral fluoresc…

cs.CL2023

A Pretrainer's Guide to Training Data: Measuring the Effects of Data Age, Domain Coverage, Quality, & Toxicity

Shayne Longpre, Gregory Yauney, Emily Reif +8

Pretraining is the preliminary and fundamental step in developing capable language models (LM). Despite this, pretraining data design is critically under-documented and often guide…

cs.CL2024

The Afterlives of Shakespeare and Company in Online Social Readership

Maria Antoniak, David Mimno, Rosamond Thalken +3

The growth of social reading platforms such as Goodreads and LibraryThing enables us to analyze reading activity at very large scale and in remarkable detail. But twenty-first cent…

cs.CL2020

Domain-Specific Lexical Grounding in Noisy Visual-Textual Documents

Gregory Yauney, Jack Hessel, David Mimno

Images can give us insights into the contextual meanings of words, but current image-text grounding approaches require detailed annotations. Such granular annotation is rare, expen…

cs.CL2026

How Reliable is Language Model Micro-Benchmarking?

Gregory Yauney, Shahzaib Saqib Warraich, Swabha Swayamdipta

Micro-benchmarking offers a solution to the often prohibitive time and cost of language model development: evaluate on a very small subset of existing benchmarks. Can these micro-b…