Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Improving Reproducibility in Evaluation through Multi-Level Annotator Modeling
Deepak Pandita, Flip Korn, Chris Welty +1
As generative AI models such as large language models (LLMs) become more pervasive, ensuring the safety, robustness, and overall trustworthiness of these systems is paramount. Howe…
cs.LG2026
How Many Ratings per Item are Necessary for Reliable Significance Testing?
Christopher Homan, Flip Korn, Deepak Pandita +1
A cornerstone of machine learning evaluation is the (often hidden) assumption that model and human responses are reliable enough to evaluate models against unitary, authoritative,…
cs.LG2025
Forest vs Tree: The Trade-off in Reproducible ML Evaluation
Deepak Pandita, Flip Korn, Chris Welty +1
Reproducibility is a cornerstone of scientific validation and of the authority it confers on its results. Reproducibility in machine learning evaluations leads to greater trust, co…