papers

Publications (32)

cs.CV2026

Evaluating Intellectual Property Guardrails of Generative Image Models: A Technical Report

Austin T. Hoag, Apostolos Modas, Yunhao Ba +9

Generative image models are capable of producing images that bear a strong resemblance to, or replicate, recognizable intellectual property (IP). In this technical report, we prese…

cs.CL2024

Not My Voice! A Taxonomy of Ethical and Safety Harms of Speech Generators

Wiebke Hutiri, Oresiti Papakyriakopoulos, Alice Xiang

The rapid and wide-scale adoption of AI to generate human speech poses a range of significant ethical and safety risks to society that need to be addressed. For example, a growing…

cs.CV2023

Men Also Do Laundry: Multi-Attribute Bias Amplification

Dora Zhao, Jerone T. A. Andrews, Alice Xiang

As computer vision systems become more widely deployed, there is increasing concern from both the research community and the public that these systems are not only reproducing but…

cs.CY2022

Regulating Facial Processing Technologies: Tensions Between Legal and Technical Considerations in the Application of Illinois BIPA

Rui-Jie Yew, Alice Xiang

Harms resulting from the development and deployment of facial processing technologies (FPT) have been met with increasing controversy. Several states and cities in the U.S. have ba…

cs.CY2023

Considerations for Ethical Speech Recognition Datasets

Orestis Papakyriakopoulos, Alice Xiang

Speech AI Technologies are largely trained on publicly available datasets or by the massive web-crawling of speech. In both cases, data acquisition focuses on minimizing collection…

cs.CV2022

From Single-Visit to Multi-Visit Image-Based Models: Single-Visit Models are Enough to Predict Obstructive Hydronephrosis

Stanley Bryan Z. Hua, Mandy Rickard, John Weaver +8

Previous work has shown the potential of deep learning to predict renal obstruction using kidney ultrasound images. However, these image-based classifiers have been trained with th…

cs.CY2020

Machine Learning Explainability for External Stakeholders

Umang Bhatt, McKane Andrus, Adrian Weller +1

As machine learning is increasingly deployed in high-stakes contexts affecting people's livelihoods, there have been growing calls to open the black box and to make machine learnin…

cs.LG2024

A Taxonomy of Challenges to Curating Fair Datasets

Dora Zhao, Morgan Klaus Scheuerman, Pooja Chitre +5

Despite extensive efforts to create fairer machine learning (ML) datasets, there remains a limited understanding of the practical aspects of dataset curation. Drawing from intervie…

cs.CV2023

Beyond Skin Tone: A Multidimensional Measure of Apparent Skin Color

William Thong, Przemyslaw Joniak, Alice Xiang

This paper strives to measure apparent skin color in computer vision, beyond a unidimensional scale on skin tone. In their seminal paper Gender Shades, Buolamwini and Gebru have sh…

cs.LG2020

Explainable Machine Learning in Deployment

Umang Bhatt, Alice Xiang, Shubham Sharma +7

Explainable machine learning offers the potential to provide stakeholders with insights into model behavior by using various methods such as feature importance scores, counterfactu…

cs.CY2019

On the Legal Compatibility of Fairness Definitions

Alice Xiang, Inioluwa Deborah Raji

Past literature has been effective in demonstrating ideological gaps in machine learning (ML) fairness definitions when considering their use in complex socio-technical systems. Ho…

cs.CY2021

"What We Can't Measure, We Can't Understand": Challenges to Demographic Data Procurement in the Pursuit of Fairness

McKane Andrus, Elena Spitzer, Jeffrey Brown +1

As calls for fair and unbiased algorithmic systems increase, so too does the number of individuals working on algorithmic fairness in industry. However, these practitioners often d…

cs.CV2024

Resampled Datasets Are Not Enough: Mitigating Societal Bias Beyond Single Attributes

Yusuke Hirota, Jerone T. A. Andrews, Dora Zhao +4

We tackle societal bias in image-text datasets by removing spurious correlations between protected groups and image attributes. Traditional methods only target labeled attributes,…

cs.CY2026

Yes, But Not Always. Generative AI Needs Nuanced Opt-in

Wiebke Hutiri, Morgan Scheuerman, Shruti Nagpal +2

This paper argues that a one-size-fits-all approach to specifying consent for the use of creative works in generative AI is insufficient. Real-world ownership and rights holder str…

cs.CY2025

TEDI: Trustworthy and Ethical Dataset Indicators to Analyze and Compare Dataset Documentation

Wiebke Hutiri, Mircea Cimpoi, Morgan Scheuerman +2

Dataset transparency is a key enabler of responsible AI, but insights into multimodal dataset attributes that impact trustworthy and ethical aspects of AI applications remain scarc…

stat.AP2015

Assessing the Potential Impact of a Nationwide Class-Based Affirmative Action System

Alice Xiang, Donald B. Rubin

We examine the possible consequences of a change in law school admissions in the United States from an affirmative action system based on race to one based on socioeconomic class.…

stat.AP2021

On the Validity of Arrest as a Proxy for Offense: Race and the Likelihood of Arrest for Violent Crimes

Riccardo Fogliato, Alice Xiang, Zachary Lipton +2

The risk of re-offense is considered in decision-making at many stages of the criminal justice system, from pre-trial, to sentencing, to parole. To aid decision makers in their ass…

cs.CY2023

Augmented Datasheets for Speech Datasets and Ethical Decision-Making

Orestis Papakyriakopoulos, Anna Seo Gyeong Choi, Jerone Andrews +5

Speech datasets are crucial for training Speech Language Technologies (SLT); however, the lack of diversity of the underlying training data can lead to serious limitations in build…

cs.SD2025

Attribution-by-design: Ensuring Inference-Time Provenance in Generative Music Systems

Fabio Morreale, Wiebke Hutiri, Joan Serrà +2

The rise of AI-generated music is diluting royalty pools and revealing structural flaws in existing remuneration frameworks, challenging the well-established artist compensation sy…

cs.CV2023

Ethical Considerations for Responsible Data Curation

Jerone T. A. Andrews, Dora Zhao, William Thong +3

Human-centric computer vision (HCCV) data curation practices often neglect privacy and bias concerns, leading to dataset retractions and unfair models. HCCV datasets constructed th…

cs.LG2022

Promises and Challenges of Causality for Ethical Machine Learning

Aida Rahmattalabi, Alice Xiang

In recent years, there has been increasing interest in causal reasoning for designing fair decision-making systems due to its compatibility with legal frameworks, interpretability…

stat.ME2023

Estimating the Likelihood of Arrest from Police Records in Presence of Unreported Crimes

Riccardo Fogliato, Arun Kumar Kuchibhotla, Zachary Lipton +3

Many important policy decisions concerning policing hinge on our understanding of how likely various criminal offenses are to result in arrests. Since many crimes are never reporte…

cs.LG2024

Efficient Bias Mitigation Without Privileged Information

Mateo Espinosa Zarlenga, Swami Sankaranarayanan, Jerone T. A. Andrews +3

Deep neural networks trained via empirical risk minimisation often exhibit significant performance disparities across groups, particularly when group and task labels are spuriously…

cs.CV2023

Flickr Africa: Examining Geo-Diversity in Large-Scale, Human-Centric Visual Data

Keziah Naggita, Julienne LaChance, Alice Xiang

Biases in large-scale image datasets are known to influence the performance of computer vision models as a function of geographic context. To investigate the limitations of standar…

cs.LG2024

Position: Measure Dataset Diversity, Don't Just Claim It

Dora Zhao, Jerone T. A. Andrews, Orestis Papakyriakopoulos +1

Machine learning (ML) datasets, often perceived as neutral, inherently encapsulate abstract and disputed social constructs. Dataset curators frequently employ value-laden terms suc…

cs.LG2025

Towards the Use of Saliency Maps for Explaining Low-Quality Electrocardiograms to End Users

Ana Lucic, Sheeraz Ahmad, Amanda Furtado Brinhosa +7

When using medical images for diagnosis, either by clinicians or artificial intelligence (AI) systems, it is important that the images are of high quality. When an image is of low…

cs.HC2021

A Multistakeholder Approach Towards Evaluating AI Transparency Mechanisms

Ana Lucic, Madhulika Srikumar, Umang Bhatt +4

Given that there are a variety of stakeholders involved in, and affected by, decisions from machine learning (ML) models, it is important to consider that different stakeholders ha…

cs.CY2026

Beyond Explainable AI (XAI): An Overdue Paradigm Shift and Post-XAI Research Directions

Saleh Afroogh, Syed Ishtiaque Ahmed, Petra Ahrweiler +46

This study provides a cross-disciplinary examination of Explainable Artificial Intelligence (XAI) approaches-focusing on deep neural networks (DNNs) and large language models (LLMs…

cs.CY2021

Uncertainty as a Form of Transparency: Measuring, Communicating, and Using Uncertainty

Umang Bhatt, Javier Antorán, Yunfeng Zhang +12

Algorithmic transparency entails exposing system properties to various stakeholders for purposes that include understanding, improving, and contesting predictions. Until now, most…

cs.CL2023

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao +448

Language models demonstrate both quantitative improvement and new qualitative capabilities with increasing scale. Despite their potentially transformative impact, these new capabil…

cs.CY2020

Affirmative Algorithms: The Legal Grounds for Fairness as Awareness

Daniel E. Ho, Alice Xiang

While there has been a flurry of research in algorithmic fairness, what is less recognized is that modern antidiscrimination law may prohibit the adoption of such techniques. We ma…

cs.CV2023

A View From Somewhere: Human-Centric Face Representations

Jerone T. A. Andrews, Przemyslaw Joniak, Alice Xiang

Few datasets contain self-identified sensitive attributes, inferring attributes risks introducing additional biases, and collecting attributes can carry legal risks. Besides, categ…