9 papers
Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge
Hanna Wallach, Meera Desai, A. Feder Cooper +17
The measurement tasks involved in evaluating generative AI (GenAI) systems lack sufficient scientific rigor, leading to what has been described as "a tangle of sloppy tests [and] a…
Examining the Expanding Role of Synthetic Data Throughout the AI Development Pipeline
Shivani Kapania, Stephanie Ballard, Alex Kessler +1
Alongside the growth of generative AI, we are witnessing a surge in the use of synthetic data across all stages of the AI development pipeline. It is now common practice for resear…
Canvil: Designerly Adaptation for LLM-Powered User Experiences
K. J. Kevin Feng, Q. Vera Liao, Ziang Xiao +3
Advancements in large language models (LLMs) are sparking a proliferation of LLM-powered user experiences (UX). In product teams, designers often craft UX to meet user needs, but i…
Fostering Appropriate Reliance on Large Language Models: The Role of Explanations, Sources, and Inconsistencies
Sunnie S. Y. Kim, Jennifer Wortman Vaughan, Q. Vera Liao +2
Large language models (LLMs) can produce erroneous responses that sound fluent and convincing, raising the risk that users will rely on these responses as if they were correct. Mit…
Supporting Industry Computing Researchers in Assessing, Articulating, and Addressing the Potential Negative Societal Impact of Their Work
Wesley Hanwen Deng, Solon Barocas, Jennifer Wortman Vaughan
Recent years have witnessed increasing calls for computing researchers to grapple with the societal impacts of their work. Tools such as impact assessments have gained prominence a…
Challenges in Human-Agent Communication
Gagan Bansal, Jennifer Wortman Vaughan, Saleema Amershi +5
Remarkable advancements in modern generative foundation models have enabled the development of sophisticated and highly capable autonomous agents that can observe their environment…