2 papers
cs.CL2023
Measuring and Improving Attentiveness to Partial Inputs with Counterfactuals
Yanai Elazar, Bhargavi Paranjape, Hao Peng +5
The inevitable appearance of spurious correlations in training datasets hurts the generalization of NLP models on unseen data. Previous work has found that datasets with paired inp…
cs.CL2023
What's In My Big Data?
Yanai Elazar, Akshita Bhagia, Ian Magnusson +10
Large text corpora are the backbone of language models. However, we have a limited understanding of the content of these corpora, including general statistics, quality, social fact…