Publications (30)
Rigor in AI: Doing Rigorous AI Work Requires a Broader, Responsible AI-Informed Conception of Rigor
Alexandra Olteanu, Su Lin Blodgett, Agathe Balayn +7
In AI research and practice, rigor remains largely understood in terms of methodological rigor -- such as whether mathematical, statistical, or computational methods are correctly…
How Different Groups Prioritize Ethical Values for Responsible AI
Maurice Jakesch, Zana Buçinca, Saleema Amershi +1
Private companies, public sector organizations, and academic groups have outlined ethical values they consider important for responsible artificial intelligence technologies. While…
"I Am the One and Only, Your Cyber BFF": Understanding the Impact of GenAI Requires Understanding the Impact of Anthropomorphic AI
Myra Cheng, Alicia DeVrio, Lisa Egede +2
Many state-of-the-art generative AI (GenAI) systems are increasingly prone to anthropomorphic behaviors, i.e., to generating outputs that are perceived to be human-like. While this…
A Taxonomy of Linguistic Expressions That Contribute To Anthropomorphism of Language Technologies
Alicia DeVrio, Myra Cheng, Lisa Egede +2
Recent attention to anthropomorphism -- the attribution of human-like qualities to non-human objects or entities -- of language technologies like LLMs has sparked renewed discussio…
The KITMUS Test: Evaluating Knowledge Integration from Multiple Sources in Natural Language Understanding Systems
Akshatha Arodi, Martin Pömsl, Kaheer Suleman +3
Many state-of-the-art natural language understanding (NLU) models are based on pretrained neural language models. These models often make inferences using information from multiple…
Responsible AI Considerations in Text Summarization Research: A Review of Current Practices
Yu Lu Liu, Meng Cao, Su Lin Blodgett +3
AI and NLP publication venues have increasingly encouraged researchers to reflect on possible ethical considerations, adverse impacts, and other responsible AI issues their work mi…
Characterizing the Demographics Behind the #BlackLivesMatter Movement
Alexandra Olteanu, Ingmar Weber, Daniel Gatica-Perez
The debates on minority issues are often dominated by or held among the concerned minority: gender equality debates have often failed to engage men, while those about race fail to…
ECBD: Evidence-Centered Benchmark Design for NLP
Yu Lu Liu, Su Lin Blodgett, Jackie Chi Kit Cheung +3
Benchmarking is seen as critical to assessing progress in NLP. However, creating a benchmark involves many design decisions (e.g., which datasets to include, which metrics to use)…
Human-Centered Responsible Artificial Intelligence: Current & Future Trends
Mohammad Tahaei, Marios Constantinides, Daniele Quercia +13
In recent years, the CHI community has seen significant growth in research on Human-Centered Responsible Artificial Intelligence. While different research communities may use diffe…
Understanding and Meeting Practitioner Needs When Measuring Representational Harms Caused by LLM-Based Systems
Emma Harvey, Emily Sheng, Su Lin Blodgett +4
The NLP research community has made publicly available numerous instruments for measuring representational harms caused by large language model (LLM)-based systems. These instrumen…
FactSheets: Increasing Trust in AI Services through Supplier's Declarations of Conformity
Matthew Arnold, Rachel K. E. Bellamy, Michael Hind +10
Accuracy is an important concern for suppliers of artificial intelligence (AI) services, but considerations beyond accuracy, such as safety (which includes fairness and explainabil…
On the Social and Technical Challenges of Web Search Autosuggestion Moderation
Timothy J. Hazen, Alexandra Olteanu, Gabriella Kazai +2
Past research shows that users benefit from systems that support them in their writing and exploration tasks. The autosuggestion feature of Web search engines is an example of such…
Proceedings of FACTS-IR 2019
Alexandra Olteanu, Jean Garcia-Gathright, Maarten de Rijke +1
The proceedings list for the program of FACTS-IR 2019, the Workshop on Fairness, Accountability, Confidentiality, Transparency, and Safety in Information Retrieval held at SIGIR 20…
Evaluating Generative AI Systems is a Social Science Measurement Challenge
Hanna Wallach, Meera Desai, Nicholas Pangakis +17
Across academia, industry, and government, there is an increasing awareness that the measurement tasks involved in evaluating generative AI (GenAI) systems are especially difficult…
"One-Size-Fits-All"? Examining Expectations around What Constitute "Fair" or "Good" NLG System Behaviors
Li Lucy, Su Lin Blodgett, Milad Shokouhi +2
Fairness-related assumptions about what constitute appropriate NLG system behaviors range from invariance, where systems are expected to behave identically for social groups, to ad…
"It was 80% me, 20% AI": Seeking Authenticity in Co-Writing with Large Language Models
Angel Hsing-Chi Hwang, Q. Vera Liao, Su Lin Blodgett +2
Given the rising proliferation and diversity of AI writing assistance tools, especially those powered by large language models (LLMs), both writers and readers may have concerns ab…
Responsible AI Research Needs Impact Statements Too
Alexandra Olteanu, Michael Ekstrand, Carlos Castillo +1
All types of research, development, and policy work can have unintended, adverse consequences - work in responsible artificial intelligence (RAI), ethical AI, or ethics in AI is no…
On Defining Erasure Harms for NLP
Yu Lu Liu, Arnav Goel, Jackie Chi Kit Cheung +3
The deployment of NLP systems has raised concerns about harms they might produce, including representational harms. Recent literature has begun to conceptualize and measure one suc…
Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge
Hanna Wallach, Meera Desai, A. Feder Cooper +17
The measurement tasks involved in evaluating generative AI (GenAI) systems lack sufficient scientific rigor, leading to what has been described as "a tangle of sloppy tests [and] a…
Can Workers Meaningfully Consent to Workplace Wellbeing Technologies?
Shreya Chowdhary, Anna Kawakami, Mary L. Gray +3
Sensing technologies deployed in the workplace can unobtrusively collect detailed data about individual activities and group interactions that are otherwise difficult to capture. A…
Overcoming Failures of Imagination in AI Infused System Development and Deployment
Margarita Boyarskaya, Alexandra Olteanu, Kate Crawford
NeurIPS 2020 requested that research paper submissions include impact statements on "potential nefarious uses and the consequences of failure." However, as researchers, practitione…
Dehumanizing Machines: Mitigating Anthropomorphic Behaviors in Text Generation Systems
Myra Cheng, Su Lin Blodgett, Alicia DeVrio +2
As text generation systems' outputs are increasingly anthropomorphic -- perceived as human-like -- scholars have also increasingly raised concerns about how such outputs can lead t…
The Effect of Extremist Violence on Hateful Speech Online
Alexandra Olteanu, Carlos Castillo, Jeremy Boy +1
User-generated content online is shaped by many factors, including endogenous elements such as platform affordances and norms, as well as exogenous elements, in particular signific…
Gaps Between Research and Practice When Measuring Representational Harms Caused by LLM-Based Systems
Emma Harvey, Emily Sheng, Su Lin Blodgett +4
To facilitate the measurement of representational harms caused by large language model (LLM)-based systems, the NLP research community has produced and made publicly available nume…
AHA!: Facilitating AI Impact Assessment by Generating Examples of Harms
Zana Buçinca, Chau Minh Pham, Maurice Jakesch +3
While demands for change and accountability for harmful AI consequences mount, foreseeing the downstream effects of deploying AI systems remains a challenging task. We developed AH…
Challenges to Evaluating the Generalization of Coreference Resolution Models: A Measurement Modeling Perspective
Ian Porada, Alexandra Olteanu, Kaheer Suleman +2
It is increasingly common to evaluate the same coreference resolution (CR) model on multiple datasets. Do these multi-dataset evaluations allow us to draw meaningful conclusions ab…
Sensing Wellbeing in the Workplace, Why and For Whom? Envisioning Impacts with Organizational Stakeholders
Anna Kawakami, Shreya Chowdhary, Shamsi T. Iqbal +4
With the heightened digitization of the workplace, alongside the rise of remote and hybrid work prompted by the pandemic, there is growing corporate interest in using passive sensi…
AI Automatons: AI Systems Intended to Imitate Humans
Alexandra Olteanu, Solon Barocas, Su Lin Blodgett +3
There is a growing proliferation of AI systems designed to mimic people's behavior, work, abilities, likenesses, or humanness -- systems we dub AI automatons. Individuals, groups,…
Deconstructing NLG Evaluation: Evaluation Practices, Assumptions, and Their Implications
Kaitlyn Zhou, Su Lin Blodgett, Adam Trischler +3
There are many ways to express similar things in text, which makes evaluating natural language generation (NLG) systems difficult. Compounding this difficulty is the need to assess…
From Use to Oversight: How Mental Models Influence User Behavior and Output in AI Writing Assistants
Shalaleh Rismani, Su Lin Blodgett, Q. Vera Liao +2
AI-based writing assistants are ubiquitous, yet little is known about how users' mental models shape their use. We examine two types of mental models -- functional or related to wh…