collaborators

11 papers

cs.CL2026

On Defining Erasure Harms for NLP

Yu Lu Liu, Arnav Goel, Jackie Chi Kit Cheung +3

The deployment of NLP systems has raised concerns about harms they might produce, including representational harms. Recent literature has begun to conceptualize and measure one suc…

cs.HC2026

Learning by Chatting? Investigating the Impact of Generative AI on Information Seeking and Learning

Shravika Mittal, Su Lin Blodgett, Q. Vera Liao

Generative AI (GenAI) tools offer increasing opportunities for augmenting human cognitive tasks. Among these tasks, information seeking is being rapidly reshaped by GenAI tools, wi…

cs.CL2026

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation

Katelyn Xiaoying Mei, Yi-Li Hsu, Minjoon Choi +5

Human evaluation plays a critical role in assessing the quality of generated text. However, the reliability and reproducibility of these evaluations depend on transparent and well-…

cs.HC2026

From Use to Oversight: How Mental Models Influence User Behavior and Output in AI Writing Assistants

Shalaleh Rismani, Su Lin Blodgett, Q. Vera Liao +2

AI-based writing assistants are ubiquitous, yet little is known about how users' mental models shape their use. We examine two types of mental models -- functional or related to wh…

cs.CY2025

Rigor in AI: Doing Rigorous AI Work Requires a Broader, Responsible AI-Informed Conception of Rigor

Alexandra Olteanu, Su Lin Blodgett, Agathe Balayn +7

In AI research and practice, rigor remains largely understood in terms of methodological rigor -- such as whether mathematical, statistical, or computational methods are correctly…

cs.CY2025

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

Hanna Wallach, Meera Desai, A. Feder Cooper +17

The measurement tasks involved in evaluating generative AI (GenAI) systems lack sufficient scientific rigor, leading to what has been described as "a tangle of sloppy tests [and] a…