5 papers
Measuring and mitigating overreliance to build human-compatible AI
Lujain Ibrahim, Katherine M. Collins, Sunnie S. Y. Kim +14
Large language models (LLMs) distinguish themselves from previous technologies by functioning as collaborative ``thought partners,'' capable of engaging more fluidly in natural lan…
PersonaTeaming: Supporting Persona-Driven Red-Teaming for Generative AI
Wesley Hanwen Deng, Mingxi Yan, Sunnie S. Y. Kim +5
Recent developments in AI safety research have called for red-teaming methods that effectively surface potential risks posed by generative AI models, with growing emphasis on how r…
Examining Human-Like Behaviors in LLMs: A Multi-Dimensional Analysis of Model Behaviors, User Factors, and System Prompts
Sunnie S. Y. Kim, Margit Bowler, Leon A Gatys
Large language models (LLMs) exhibit a wide range of human-like behaviors, from expressing thoughts and emotions, to engaging in relationship-building with users, to refusing reque…
Understanding Annotator Safety Policy with Interpretability
Alex Oesterling, Donghao Ren, Yannick Assogba +4
Safety policies define what constitutes safe and unsafe AI outputs, guiding data annotation and model development. However, annotation disagreement is pervasive and can stem from m…
Fostering Appropriate Reliance on Large Language Models: The Role of Explanations, Sources, and Inconsistencies
Sunnie S. Y. Kim, Jennifer Wortman Vaughan, Q. Vera Liao +2
Large language models (LLMs) can produce erroneous responses that sound fluent and convincing, raising the risk that users will rely on these responses as if they were correct. Mit…