papers

Publications (30)

cs.CY2025

Rigor in AI: Doing Rigorous AI Work Requires a Broader, Responsible AI-Informed Conception of Rigor

Alexandra Olteanu, Su Lin Blodgett, Agathe Balayn +7

In AI research and practice, rigor remains largely understood in terms of methodological rigor -- such as whether mathematical, statistical, or computational methods are correctly…

cs.HC2022

How Different Groups Prioritize Ethical Values for Responsible AI

Maurice Jakesch, Zana Buçinca, Saleema Amershi +1

Private companies, public sector organizations, and academic groups have outlined ethical values they consider important for responsible artificial intelligence technologies. While…

cs.CY2024

"I Am the One and Only, Your Cyber BFF": Understanding the Impact of GenAI Requires Understanding the Impact of Anthropomorphic AI

Myra Cheng, Alicia DeVrio, Lisa Egede +2

Many state-of-the-art generative AI (GenAI) systems are increasingly prone to anthropomorphic behaviors, i.e., to generating outputs that are perceived to be human-like. While this…

cs.HC2025

A Taxonomy of Linguistic Expressions That Contribute To Anthropomorphism of Language Technologies

Alicia DeVrio, Myra Cheng, Lisa Egede +2

Recent attention to anthropomorphism -- the attribution of human-like qualities to non-human objects or entities -- of language technologies like LLMs has sparked renewed discussio…

cs.CL2023

The KITMUS Test: Evaluating Knowledge Integration from Multiple Sources in Natural Language Understanding Systems

Akshatha Arodi, Martin Pömsl, Kaheer Suleman +3

Many state-of-the-art natural language understanding (NLU) models are based on pretrained neural language models. These models often make inferences using information from multiple…

cs.CL2023

Responsible AI Considerations in Text Summarization Research: A Review of Current Practices

Yu Lu Liu, Meng Cao, Su Lin Blodgett +3

AI and NLP publication venues have increasingly encouraged researchers to reflect on possible ethical considerations, adverse impacts, and other responsible AI issues their work mi…

cs.SI2015

Characterizing the Demographics Behind the #BlackLivesMatter Movement

Alexandra Olteanu, Ingmar Weber, Daniel Gatica-Perez

The debates on minority issues are often dominated by or held among the concerned minority: gender equality debates have often failed to engage men, while those about race fail to…

cs.CL2024

ECBD: Evidence-Centered Benchmark Design for NLP

Yu Lu Liu, Su Lin Blodgett, Jackie Chi Kit Cheung +3

Benchmarking is seen as critical to assessing progress in NLP. However, creating a benchmark involves many design decisions (e.g., which datasets to include, which metrics to use)…

cs.HC2023

Human-Centered Responsible Artificial Intelligence: Current & Future Trends

Mohammad Tahaei, Marios Constantinides, Daniele Quercia +13

In recent years, the CHI community has seen significant growth in research on Human-Centered Responsible Artificial Intelligence. While different research communities may use diffe…

cs.CY2025

Understanding and Meeting Practitioner Needs When Measuring Representational Harms Caused by LLM-Based Systems

Emma Harvey, Emily Sheng, Su Lin Blodgett +4

The NLP research community has made publicly available numerous instruments for measuring representational harms caused by large language model (LLM)-based systems. These instrumen…

cs.CY2019

FactSheets: Increasing Trust in AI Services through Supplier's Declarations of Conformity

Matthew Arnold, Rachel K. E. Bellamy, Michael Hind +10

Accuracy is an important concern for suppliers of artificial intelligence (AI) services, but considerations beyond accuracy, such as safety (which includes fairness and explainabil…

cs.CY2020

On the Social and Technical Challenges of Web Search Autosuggestion Moderation

Timothy J. Hazen, Alexandra Olteanu, Gabriella Kazai +2

Past research shows that users benefit from systems that support them in their writing and exploration tasks. The autosuggestion feature of Web search engines is an example of such…

cs.IR2019

Proceedings of FACTS-IR 2019

Alexandra Olteanu, Jean Garcia-Gathright, Maarten de Rijke +1

The proceedings list for the program of FACTS-IR 2019, the Workshop on Fairness, Accountability, Confidentiality, Transparency, and Safety in Information Retrieval held at SIGIR 20…

cs.CY2024

Evaluating Generative AI Systems is a Social Science Measurement Challenge

Hanna Wallach, Meera Desai, Nicholas Pangakis +17

Across academia, industry, and government, there is an increasing awareness that the measurement tasks involved in evaluating generative AI (GenAI) systems are especially difficult…

cs.CL2024

"One-Size-Fits-All"? Examining Expectations around What Constitute "Fair" or "Good" NLG System Behaviors

Li Lucy, Su Lin Blodgett, Milad Shokouhi +2

Fairness-related assumptions about what constitute appropriate NLG system behaviors range from invariance, where systems are expected to behave identically for social groups, to ad…

cs.HC2024

"It was 80% me, 20% AI": Seeking Authenticity in Co-Writing with Large Language Models

Angel Hsing-Chi Hwang, Q. Vera Liao, Su Lin Blodgett +2

Given the rising proliferation and diversity of AI writing assistance tools, especially those powered by large language models (LLMs), both writers and readers may have concerns ab…

cs.AI2023

Responsible AI Research Needs Impact Statements Too

Alexandra Olteanu, Michael Ekstrand, Carlos Castillo +1

All types of research, development, and policy work can have unintended, adverse consequences - work in responsible artificial intelligence (RAI), ethical AI, or ethics in AI is no…

cs.CL2026

On Defining Erasure Harms for NLP

Yu Lu Liu, Arnav Goel, Jackie Chi Kit Cheung +3

The deployment of NLP systems has raised concerns about harms they might produce, including representational harms. Recent literature has begun to conceptualize and measure one suc…

cs.CY2025

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

Hanna Wallach, Meera Desai, A. Feder Cooper +17

The measurement tasks involved in evaluating generative AI (GenAI) systems lack sufficient scientific rigor, leading to what has been described as "a tangle of sloppy tests [and] a…

cs.HC2023

Can Workers Meaningfully Consent to Workplace Wellbeing Technologies?

Shreya Chowdhary, Anna Kawakami, Mary L. Gray +3

Sensing technologies deployed in the workplace can unobtrusively collect detailed data about individual activities and group interactions that are otherwise difficult to capture. A…

cs.CY2020

Overcoming Failures of Imagination in AI Infused System Development and Deployment

Margarita Boyarskaya, Alexandra Olteanu, Kate Crawford

NeurIPS 2020 requested that research paper submissions include impact statements on "potential nefarious uses and the consequences of failure." However, as researchers, practitione…

cs.CL2025

Dehumanizing Machines: Mitigating Anthropomorphic Behaviors in Text Generation Systems

Myra Cheng, Su Lin Blodgett, Alicia DeVrio +2

As text generation systems' outputs are increasingly anthropomorphic -- perceived as human-like -- scholars have also increasingly raised concerns about how such outputs can lead t…

cs.SI2018

The Effect of Extremist Violence on Hateful Speech Online

Alexandra Olteanu, Carlos Castillo, Jeremy Boy +1

User-generated content online is shaped by many factors, including endogenous elements such as platform affordances and norms, as well as exogenous elements, in particular signific…

cs.CY2024

Gaps Between Research and Practice When Measuring Representational Harms Caused by LLM-Based Systems

Emma Harvey, Emily Sheng, Su Lin Blodgett +4

To facilitate the measurement of representational harms caused by large language model (LLM)-based systems, the NLP research community has produced and made publicly available nume…

cs.HC2023

AHA!: Facilitating AI Impact Assessment by Generating Examples of Harms

Zana Buçinca, Chau Minh Pham, Maurice Jakesch +3

While demands for change and accountability for harmful AI consequences mount, foreseeing the downstream effects of deploying AI systems remains a challenging task. We developed AH…

cs.CL2024

Challenges to Evaluating the Generalization of Coreference Resolution Models: A Measurement Modeling Perspective

Ian Porada, Alexandra Olteanu, Kaheer Suleman +2

It is increasingly common to evaluate the same coreference resolution (CR) model on multiple datasets. Do these multi-dataset evaluations allow us to draw meaningful conclusions ab…

cs.HC2023

Sensing Wellbeing in the Workplace, Why and For Whom? Envisioning Impacts with Organizational Stakeholders

Anna Kawakami, Shreya Chowdhary, Shamsi T. Iqbal +4

With the heightened digitization of the workplace, alongside the rise of remote and hybrid work prompted by the pandemic, there is growing corporate interest in using passive sensi…

cs.CY2025

AI Automatons: AI Systems Intended to Imitate Humans

Alexandra Olteanu, Solon Barocas, Su Lin Blodgett +3

There is a growing proliferation of AI systems designed to mimic people's behavior, work, abilities, likenesses, or humanness -- systems we dub AI automatons. Individuals, groups,…

cs.CL2022

Deconstructing NLG Evaluation: Evaluation Practices, Assumptions, and Their Implications

Kaitlyn Zhou, Su Lin Blodgett, Adam Trischler +3

There are many ways to express similar things in text, which makes evaluating natural language generation (NLG) systems difficult. Compounding this difficulty is the need to assess…

cs.HC2026

From Use to Oversight: How Mental Models Influence User Behavior and Output in AI Writing Assistants

Shalaleh Rismani, Su Lin Blodgett, Q. Vera Liao +2

AI-based writing assistants are ubiquitous, yet little is known about how users' mental models shape their use. We examine two types of mental models -- functional or related to wh…