2 papers
cs.LG2026
NodeSynth: Socially Aligned Synthetic Data for AI Evaluation
Qazi Mamunur Rashid, Xuan Yang, Zhengzhe Yang +5
Recent advancements in generative AI facilitate large-scale synthetic data generation for model evaluation. However, without targeted approaches, these datasets often lack the soci…
cs.CY2024
A Toolbox for Surfacing Health Equity Harms and Biases in Large Language Models
Stephen R. Pfohl, Heather Cole-Lewis, Rory Sayres +27
Large language models (LLMs) hold promise to serve complex health information needs but also have the potential to introduce harm and exacerbate health disparities. Reliably evalua…