51 citations · 53 across the 12 of their papers we have counts for
3 papers · 1 filter
Quantifying the Statistical Effect of Rubric Modifications on Human-Autorater Agreement
Jessica Huynh, Alfredo Gomez, Athiya Deviyani +3
Autoraters, also referred to as LLM-as-judges, are increasingly used for evaluation and automated content moderation. However, there is limited statistical analysis of how modifica…
Cultural Perspectives and Expectations for Generative AI: A Global Survey Approach
Erin van Liemt, Renee Shelby, Andrew Smart +5
There is a lack of empirical evidence about global attitudes around whether and how GenAI should represent cultures. This paper assesses understandings and beliefs about culture as…
Debiasing Text Safety Classifiers through a Fairness-Aware Ensemble
Olivia Sturman, Aparna Joshi, Bhaktipriya Radharapu +2
Increasing use of large language models (LLMs) demand performant guardrails to ensure the safety of inputs and outputs of LLMs. When these safeguards are trained on imbalanced data…