2 papers
cs.LG2026
Improving Reproducibility in Evaluation through Multi-Level Annotator Modeling
Deepak Pandita, Flip Korn, Chris Welty +1
As generative AI models such as large language models (LLMs) become more pervasive, ensuring the safety, robustness, and overall trustworthiness of these systems is paramount. Howe…
cs.LG2025
Forest vs Tree: The Trade-off in Reproducible ML Evaluation
Deepak Pandita, Flip Korn, Chris Welty +1
Reproducibility is a cornerstone of scientific validation and of the authority it confers on its results. Reproducibility in machine learning evaluations leads to greater trust, co…