9 papers
RamseyGadgets: A Graph Construction Dataset for LLMs
Zohair Raza Hassan, Deepak Pandita
Constructing special graphs is an important task within graph theory and computer science. Many popular graph constructions are the result of a comprehensive exploration of relevan…
Pluralis v0.1: Towards a Multicultural, Multimodal, Multilingual Benchmark for AI Risk and Reliability
Alicia Parrish, Rajat Shinde, Sanket Badhe +57
Current AI safety evaluation and benchmarking frameworks predominantly rely on Western-centric culture-agnostic defaults that mask critical regional laws, socio-linguistic nuances,…
Improving Reproducibility in Evaluation through Multi-Level Annotator Modeling
Deepak Pandita, Flip Korn, Chris Welty +1
As generative AI models such as large language models (LLMs) become more pervasive, ensuring the safety, robustness, and overall trustworthiness of these systems is paramount. Howe…
Learning Who Disagrees: Demographic Importance Weighting for Modeling Annotator Distributions with DiADEM
Samay U. Shetty, Tharindu Cyril Weerasooriya, Deepak Pandita +1
When humans label subjective content, they disagree, and that disagreement is not noise. It reflects genuine differences in perspective shaped by annotators' social identities and…
How Many Ratings per Item are Necessary for Reliable Significance Testing?
Christopher Homan, Flip Korn, Deepak Pandita +1
A cornerstone of machine learning evaluation is the (often hidden) assumption that model and human responses are reliable enough to evaluate models against unitary, authoritative,…
Forest vs Tree: The Trade-off in Reproducible ML Evaluation
Deepak Pandita, Flip Korn, Chris Welty +1
Reproducibility is a cornerstone of scientific validation and of the authority it confers on its results. Reproducibility in machine learning evaluations leads to greater trust, co…