5 papers
Improving Reproducibility in Evaluation through Multi-Level Annotator Modeling
Deepak Pandita, Flip Korn, Chris Welty +1
As generative AI models such as large language models (LLMs) become more pervasive, ensuring the safety, robustness, and overall trustworthiness of these systems is paramount. Howe…
Computer Science Conferences Should Require Nonrepudiable Experimental Results
Mamadou K. Keita, Christopher Homan
This position paper argues that computer science conferences should require tamper-evident, nonrepudiable attestations of experimental results. We name the underlying problem exper…
Learning Who Disagrees: Demographic Importance Weighting for Modeling Annotator Distributions with DiADEM
Samay U. Shetty, Tharindu Cyril Weerasooriya, Deepak Pandita +1
When humans label subjective content, they disagree, and that disagreement is not noise. It reflects genuine differences in perspective shaped by annotators' social identities and…
How Many Ratings per Item are Necessary for Reliable Significance Testing?
Christopher Homan, Flip Korn, Deepak Pandita +1
A cornerstone of machine learning evaluation is the (often hidden) assumption that model and human responses are reliable enough to evaluate models against unitary, authoritative,…
Forest vs Tree: The Trade-off in Reproducible ML Evaluation
Deepak Pandita, Flip Korn, Chris Welty +1
Reproducibility is a cornerstone of scientific validation and of the authority it confers on its results. Reproducibility in machine learning evaluations leads to greater trust, co…