9 papers
Beyond "I Can't Help With That": How Child Safety Experts Evaluate AI Chatbot Safety
Hannah Cha, Neha Shukla, Solon Barocas +3
Youth increasingly turn to AI chatbots for social and emotional support, raising concerns about how these systems respond, especially in high-stakes situations. However, existing c…
Psychological Influences of Conversational AI: Research and Design Directions for Reducing Harm and Promoting Well-Being
Jina Suh, Mihaela Vorvoreanu, Forough Poursabzi-Sangdeh +5
As conversational AI systems become increasingly integrated into daily life, their potential effects on user well-being require ongoing attention. While consumer-facing generalist…
Pluralis v0.1: Towards a Multicultural, Multimodal, Multilingual Benchmark for AI Risk and Reliability
Alicia Parrish, Rajat Shinde, Sanket Badhe +57
Current AI safety evaluation and benchmarking frameworks predominantly rely on Western-centric culture-agnostic defaults that mask critical regional laws, socio-linguistic nuances,…
Learning to Attack and Defend: Adaptive Red Teaming of Language Models via GRPO
Blake Bullwinkel, Eugenia Kim, Amanda Minnich +1
AI red teaming must continually adapt to evolving attackers and defenders. Reinforcement learning offers a promising approach to discovering novel attacks, and co-training methods…
DisaBench: A Participatory Evaluation Framework for Disability Harms in Language Models
Eugenia Kim, Ioana Tanase, Christina Mallon
General-purpose safety benchmarks for large language models do not adequately evaluate disability-related harms. We introduce DisaBench: a taxonomy of twelve disability harm catego…
XL-SafetyBench: A Country-Grounded Cross-Cultural Benchmark for LLM Safety and Cultural Sensitivity
Dasol Choi, Eugenia Kim, Jaewon Noh +14
Current LLM safety benchmarks are predominantly English-centric and often rely on translation, failing to capture country-specific harms. Moreover, they rarely evaluate a model's a…