Publications (32)
Evaluating Intellectual Property Guardrails of Generative Image Models: A Technical Report
Austin T. Hoag, Apostolos Modas, Yunhao Ba +9
Generative image models are capable of producing images that bear a strong resemblance to, or replicate, recognizable intellectual property (IP). In this technical report, we prese…
Not My Voice! A Taxonomy of Ethical and Safety Harms of Speech Generators
Wiebke Hutiri, Oresiti Papakyriakopoulos, Alice Xiang
The rapid and wide-scale adoption of AI to generate human speech poses a range of significant ethical and safety risks to society that need to be addressed. For example, a growing…
Men Also Do Laundry: Multi-Attribute Bias Amplification
Dora Zhao, Jerone T. A. Andrews, Alice Xiang
As computer vision systems become more widely deployed, there is increasing concern from both the research community and the public that these systems are not only reproducing but…
Regulating Facial Processing Technologies: Tensions Between Legal and Technical Considerations in the Application of Illinois BIPA
Rui-Jie Yew, Alice Xiang
Harms resulting from the development and deployment of facial processing technologies (FPT) have been met with increasing controversy. Several states and cities in the U.S. have ba…
Considerations for Ethical Speech Recognition Datasets
Orestis Papakyriakopoulos, Alice Xiang
Speech AI Technologies are largely trained on publicly available datasets or by the massive web-crawling of speech. In both cases, data acquisition focuses on minimizing collection…
From Single-Visit to Multi-Visit Image-Based Models: Single-Visit Models are Enough to Predict Obstructive Hydronephrosis
Stanley Bryan Z. Hua, Mandy Rickard, John Weaver +8
Previous work has shown the potential of deep learning to predict renal obstruction using kidney ultrasound images. However, these image-based classifiers have been trained with th…
Machine Learning Explainability for External Stakeholders
Umang Bhatt, McKane Andrus, Adrian Weller +1
As machine learning is increasingly deployed in high-stakes contexts affecting people's livelihoods, there have been growing calls to open the black box and to make machine learnin…
A Taxonomy of Challenges to Curating Fair Datasets
Dora Zhao, Morgan Klaus Scheuerman, Pooja Chitre +5
Despite extensive efforts to create fairer machine learning (ML) datasets, there remains a limited understanding of the practical aspects of dataset curation. Drawing from intervie…
Beyond Skin Tone: A Multidimensional Measure of Apparent Skin Color
William Thong, Przemyslaw Joniak, Alice Xiang
This paper strives to measure apparent skin color in computer vision, beyond a unidimensional scale on skin tone. In their seminal paper Gender Shades, Buolamwini and Gebru have sh…
Explainable Machine Learning in Deployment
Umang Bhatt, Alice Xiang, Shubham Sharma +7
Explainable machine learning offers the potential to provide stakeholders with insights into model behavior by using various methods such as feature importance scores, counterfactu…
On the Legal Compatibility of Fairness Definitions
Alice Xiang, Inioluwa Deborah Raji
Past literature has been effective in demonstrating ideological gaps in machine learning (ML) fairness definitions when considering their use in complex socio-technical systems. Ho…
"What We Can't Measure, We Can't Understand": Challenges to Demographic Data Procurement in the Pursuit of Fairness
McKane Andrus, Elena Spitzer, Jeffrey Brown +1
As calls for fair and unbiased algorithmic systems increase, so too does the number of individuals working on algorithmic fairness in industry. However, these practitioners often d…
Resampled Datasets Are Not Enough: Mitigating Societal Bias Beyond Single Attributes
Yusuke Hirota, Jerone T. A. Andrews, Dora Zhao +4
We tackle societal bias in image-text datasets by removing spurious correlations between protected groups and image attributes. Traditional methods only target labeled attributes,…
Yes, But Not Always. Generative AI Needs Nuanced Opt-in
Wiebke Hutiri, Morgan Scheuerman, Shruti Nagpal +2
This paper argues that a one-size-fits-all approach to specifying consent for the use of creative works in generative AI is insufficient. Real-world ownership and rights holder str…
TEDI: Trustworthy and Ethical Dataset Indicators to Analyze and Compare Dataset Documentation
Wiebke Hutiri, Mircea Cimpoi, Morgan Scheuerman +2
Dataset transparency is a key enabler of responsible AI, but insights into multimodal dataset attributes that impact trustworthy and ethical aspects of AI applications remain scarc…
Assessing the Potential Impact of a Nationwide Class-Based Affirmative Action System
Alice Xiang, Donald B. Rubin
We examine the possible consequences of a change in law school admissions in the United States from an affirmative action system based on race to one based on socioeconomic class.…
On the Validity of Arrest as a Proxy for Offense: Race and the Likelihood of Arrest for Violent Crimes
Riccardo Fogliato, Alice Xiang, Zachary Lipton +2
The risk of re-offense is considered in decision-making at many stages of the criminal justice system, from pre-trial, to sentencing, to parole. To aid decision makers in their ass…
Augmented Datasheets for Speech Datasets and Ethical Decision-Making
Orestis Papakyriakopoulos, Anna Seo Gyeong Choi, Jerone Andrews +5
Speech datasets are crucial for training Speech Language Technologies (SLT); however, the lack of diversity of the underlying training data can lead to serious limitations in build…
Attribution-by-design: Ensuring Inference-Time Provenance in Generative Music Systems
Fabio Morreale, Wiebke Hutiri, Joan Serrà +2
The rise of AI-generated music is diluting royalty pools and revealing structural flaws in existing remuneration frameworks, challenging the well-established artist compensation sy…
Ethical Considerations for Responsible Data Curation
Jerone T. A. Andrews, Dora Zhao, William Thong +3
Human-centric computer vision (HCCV) data curation practices often neglect privacy and bias concerns, leading to dataset retractions and unfair models. HCCV datasets constructed th…
Promises and Challenges of Causality for Ethical Machine Learning
Aida Rahmattalabi, Alice Xiang
In recent years, there has been increasing interest in causal reasoning for designing fair decision-making systems due to its compatibility with legal frameworks, interpretability…
Estimating the Likelihood of Arrest from Police Records in Presence of Unreported Crimes
Riccardo Fogliato, Arun Kumar Kuchibhotla, Zachary Lipton +3
Many important policy decisions concerning policing hinge on our understanding of how likely various criminal offenses are to result in arrests. Since many crimes are never reporte…
Efficient Bias Mitigation Without Privileged Information
Mateo Espinosa Zarlenga, Swami Sankaranarayanan, Jerone T. A. Andrews +3
Deep neural networks trained via empirical risk minimisation often exhibit significant performance disparities across groups, particularly when group and task labels are spuriously…
Flickr Africa: Examining Geo-Diversity in Large-Scale, Human-Centric Visual Data
Keziah Naggita, Julienne LaChance, Alice Xiang
Biases in large-scale image datasets are known to influence the performance of computer vision models as a function of geographic context. To investigate the limitations of standar…
Position: Measure Dataset Diversity, Don't Just Claim It
Dora Zhao, Jerone T. A. Andrews, Orestis Papakyriakopoulos +1
Machine learning (ML) datasets, often perceived as neutral, inherently encapsulate abstract and disputed social constructs. Dataset curators frequently employ value-laden terms suc…
Towards the Use of Saliency Maps for Explaining Low-Quality Electrocardiograms to End Users
Ana Lucic, Sheeraz Ahmad, Amanda Furtado Brinhosa +7
When using medical images for diagnosis, either by clinicians or artificial intelligence (AI) systems, it is important that the images are of high quality. When an image is of low…
A Multistakeholder Approach Towards Evaluating AI Transparency Mechanisms
Ana Lucic, Madhulika Srikumar, Umang Bhatt +4
Given that there are a variety of stakeholders involved in, and affected by, decisions from machine learning (ML) models, it is important to consider that different stakeholders ha…
Beyond Explainable AI (XAI): An Overdue Paradigm Shift and Post-XAI Research Directions
Saleh Afroogh, Syed Ishtiaque Ahmed, Petra Ahrweiler +46
This study provides a cross-disciplinary examination of Explainable Artificial Intelligence (XAI) approaches-focusing on deep neural networks (DNNs) and large language models (LLMs…
Uncertainty as a Form of Transparency: Measuring, Communicating, and Using Uncertainty
Umang Bhatt, Javier Antorán, Yunfeng Zhang +12
Algorithmic transparency entails exposing system properties to various stakeholders for purposes that include understanding, improving, and contesting predictions. Until now, most…
Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao +448
Language models demonstrate both quantitative improvement and new qualitative capabilities with increasing scale. Despite their potentially transformative impact, these new capabil…
Affirmative Algorithms: The Legal Grounds for Fairness as Awareness
Daniel E. Ho, Alice Xiang
While there has been a flurry of research in algorithmic fairness, what is less recognized is that modern antidiscrimination law may prohibit the adoption of such techniques. We ma…
A View From Somewhere: Human-Centric Face Representations
Jerone T. A. Andrews, Przemyslaw Joniak, Alice Xiang
Few datasets contain self-identified sensitive attributes, inferring attributes risks introducing additional biases, and collecting attributes can carry legal risks. Besides, categ…