3 papers
stat.ME2026
From Ground Truth to Measurement: A Statistical Framework for Human Labeling
Robert Chew, Stephanie Eckman, Christoph Kern +1
Supervised machine learning assumes that labeled data provide accurate measurements of the concepts models are meant to learn. Yet in practice, human labeling introduces systematic…
stat.ME2025
Aligning NLP Models with Target Population Perspectives using PAIR: Population-Aligned Instance Replication
Stephanie Eckman, Bolei Ma, Christoph Kern +3
Models trained on crowdsourced annotations may not reflect population views, if those who work as annotators do not represent the broader population. In this paper, we propose PAIR…
stat.ML2023
Annotation Sensitivity: Training Data Collection Methods Affect Model Performance
Christoph Kern, Stephanie Eckman, Jacob Beck +3
When training data are collected from human annotators, the design of the annotation instrument, the instructions given to annotators, the characteristics of the annotators, and th…