Showing cs.CYShow all
2 papers · 1 filter
cs.CY2026
Decoding Safety Feedback from Diverse Raters: A Data-driven Lens on Responsiveness to Severity
Pushkar Mishra, Charvi Rastogi, Stephen R. Pfohl +9
Ensuring the safety of Generative AI requires a nuanced understanding of pluralistic viewpoints. In this paper, we introduce a novel data-driven approach for analyzing ordinal safe…
cs.CY2024
A Toolbox for Surfacing Health Equity Harms and Biases in Large Language Models
Stephen R. Pfohl, Heather Cole-Lewis, Rory Sayres +27
Large language models (LLMs) hold promise to serve complex health information needs but also have the potential to introduce harm and exacerbate health disparities. Reliably evalua…