papers

Publications (30)

cs.CY2025

A Framework for Auditing Chatbots for Dialect-Based Quality-of-Service Harms

Emma Harvey, Rene F. Kizilcec, Allison Koenecke

Increasingly, individuals who engage in online activities are expected to interact with large language model (LLM)-based chatbots. Prior work has shown that LLMs can display dialec…

cs.CL2019

Learning Twitter User Sentiments on Climate Change with Limited Labeled Data

Allison Koenecke, Jordi Feliu-FabÃ

While it is well-documented that climate change accepters and deniers have become increasingly polarized in the United States over time, there has been no large-scale examination o…

cs.LG2019

Curriculum Learning in Deep Neural Networks for Financial Forecasting

Allison Koenecke, Amita Gajewar

For any financial organization, computing accurate quarterly forecasts for various products is one of the most critical operations. As the granularity at which forecasts are needed…

cs.CY2025

Bias Delayed is Bias Denied? Assessing the Effect of Reporting Delays on Disparity Assessments

Jennah Gosciak, Aparna Balagopalan, Derek Ouyang +3

Conducting disparity assessments at regular time intervals is critical for surfacing potential biases in decision-making and improving outcomes across demographic groups. Because d…

stat.AP2023

Potential for allocative harm in an environmental justice data tool

Benjamin Q. Huynh, Elizabeth T. Chin, Allison Koenecke +4

Neighborhood-level screening algorithms are increasingly being deployed to inform policy decisions. We evaluate one such algorithm, CalEnviroScreen - designed to promote environmen…

cs.CY2023

Popular Support for Balancing Equity and Efficiency in Resource Allocation: A Case Study in Online Advertising to Increase Welfare Program Awareness

Allison Koenecke, Eric Giannella, Robb Willer +1

Algorithmically optimizing the provision of limited resources is commonplace across domains from healthcare to lending. Optimization can lead to efficient resource allocation, but,…

q-bio.TO2021

Alpha-1 adrenergic receptor antagonists to prevent hyperinflammation and death from lower respiratory tract infection

Allison Koenecke, Michael Powell, Ruoxuan Xiong +22

In severe viral pneumonia, including Coronavirus disease 2019 (COVID-19), the viral replication phase is often followed by hyperinflammation, which can lead to acute respiratory di…

cs.CL2024

Careless Whisper: Speech-to-Text Hallucination Harms

Allison Koenecke, Anna Seo Gyeong Choi, Katelyn X. Mei +2

Speech-to-text services aim to transcribe input audio as accurately as possible. They increasingly play a role in everyday life, for example in personal voice assistants or in cust…

cs.CL2025

Analyzing Dialectical Biases in LLMs for Knowledge and Reasoning Benchmarks

Eileen Pan, Anna Seo Gyeong Choi, Maartje ter Hoeve +2

Large language models (LLMs) are ubiquitous in modern day natural language processing. However, previous work has shown degraded LLM performance for under-represented English diale…

cs.LG2023

Federated Causal Inference in Heterogeneous Observational Data

Ruoxuan Xiong, Allison Koenecke, Michael Powell +3

We are interested in estimating the effect of a treatment applied to individuals at multiple sites, where data is stored locally for each site. Due to privacy constraints, individu…

econ.GN2019

A Game Theoretic Setting of Capitation Versus Fee-For-Service Payment Systems

Allison Koenecke

We aim to determine whether a game-theoretic model between an insurer and a healthcare practice yields a predictive equilibrium that incentivizes either player to deviate from a fe…

cs.CL2025

Characterizing Bias: Benchmarking Large Language Models in Simplified versus Traditional Chinese

Hanjia Lyu, Jiebo Luo, Jian Kang +1

While the capabilities of Large Language Models (LLMs) have been studied in both Simplified and Traditional Chinese, it is yet unclear whether LLMs exhibit differential performance…

cs.HC2026

LLMs in social services: How does chatbot accuracy affect human accuracy?

Jennah Gosciak, Eric Giannella, Zhaowen Guo +2

Social service programs like the Supplemental Nutrition Assistance Program (SNAP, or food stamps) have eligibility rules that can be challenging to understand. For nonprofit casewo…

cs.CY2023

Augmented Datasheets for Speech Datasets and Ethical Decision-Making

Orestis Papakyriakopoulos, Anna Seo Gyeong Choi, Jerone Andrews +5

Speech datasets are crucial for training Speech Language Technologies (SLT); however, the lack of diversity of the underlying training data can lead to serious limitations in build…

cs.HC2022

Trucks Don't Mean Trump: Diagnosing Human Error in Image Analysis

J. D. Zamfirescu-Pereira, Jerry Chen, Emily Wen +3

Algorithms provide powerful tools for detecting and dissecting human bias and error. Here, we develop machine learning methods to to analyze how humans err in a particular high-sta…

cs.CL2024

Automate or Assist? The Role of Computational Models in Identifying Gendered Discourse in US Capital Trial Transcripts

Andrea W Wen-Yi, Kathryn Adamson, Nathalie Greenfield +4

The language used by US courtroom actors in criminal trials has long been studied for biases. However, systematic studies for bias in high-stakes court trials have been difficult,…

cs.CV2025

SPHERE: Unveiling Spatial Blind Spots in Vision-Language Models Through Hierarchical Evaluation

Wenyu Zhang, Wei En Ng, Lixin Ma +5

Current vision-language models may grasp basic spatial cues and simple directions (e.g. left, right, front, back), but struggle with the multi-dimensional spatial reasoning necessa…

econ.GN2020

Synthetic Data Generation for Economists

Allison Koenecke, Hal Varian

As more tech companies engage in rigorous economic analyses, we are confronted with a data problem: in-house papers cannot be replicated due to use of sensitive, proprietary, or pr…

cs.CY2026

Introducing AI to an Online Petition Platform Changed Outputs but not Outcomes

Isabel Corpus, Eric Gilbert, Allison Koenecke +1

The rapid integration of AI writing tools into online platforms raises critical questions about their impact on content production and outcomes. We leverage a unique natural experi…

cs.IR2023

Auditing Cross-Cultural Consistency of Human-Annotated Labels for Recommendation Systems

Rock Yuren Pang, Jack Cenatempo, Franklyn Graham +5

Recommendation systems increasingly depend on massive human-labeled datasets; however, the human annotators hired to generate these labels increasingly come from homogeneous backgr…

cs.CY2026

Addressing Pitfalls in Auditing Practices of Automatic Speech Recognition Technologies: A Case Study of People with Aphasia

Katelyn Xiaoying Mei, Anna Seo Gyeong Choi, Hilke Schellmann +2

Automatic Speech Recognition (ASR) systems' growing use warrants robust auditing approaches to ensure equitable transcription quality, especially for people with speech disorders l…

cs.CL2025

Tasks and Roles in Legal AI: Data Curation, Annotation, and Verification

Allison Koenecke, Jed Stiglitz, David Mimno +1

The application of AI tools to the legal field feels natural: large legal document collections could be used with specialized AI to improve workflow efficiency for lawyers and amel…

quant-ph2013

Recovering the Period in Shor's Algorithm with Gauss' Algorithm for Lattice Basis Reduction

Allison Koenecke, Pawel Wocjan

Shor's algorithm contains a classical post-processing part for which we aim to create an efficient, understandable method aside from continued fractions. Let r be an unknown positi…

cs.CY2026

A Critical Pragmatism Approach for Algorithmic Fairness: Lessons from Urban Planning Theory

Jennah Gosciak, Karen Levy, Allison Koenecke

As data scientists grapple with increasingly complex ethical decisions in machine learning (ML) and data science, the field of algorithmic fairness has offered multiple solutions,…

cs.AI2025

Operationalizing Pluralistic Values in Large Language Model Alignment Reveals Trade-offs in Safety, Inclusivity, and Model Behavior

Dalia Ali, Dora Zhao, Allison Koenecke +1

Although large language models (LLMs) are increasingly trained using human feedback for safety and alignment with human values, alignment decisions often overlook human social dive…

cs.CY2025

"Don't Forget the Teachers": Towards an Educator-Centered Understanding of Harms from Large Language Models in Education

Emma Harvey, Allison Koenecke, Rene F. Kizilcec

Education technologies (edtech) are increasingly incorporating new features built on large language models (LLMs), with the goals of enriching the processes of teaching and learnin…

cs.CY2026

Into the Unknown: Accounting for Missing Demographic Data when Mitigating Ad Delivery Skew

Isabel Corpus, Allison Koenecke

Online advertising platforms use algorithmic systems to power the process of matching ads to users, termed ad delivery. Prior audits have demonstrated that ad delivery can be skewe…

cs.CY2026

Scrutinizing Index-Based Risk Assessments: A Case Study in NYC Decision-making for Heat Emergency Management

Jennah Gosciak, Luke Boyce, Angelina Wang +1

Cities are increasingly turning to large-scale data analysis and machine learning to make consequential decisions. While the algorithmic fairness community has focused on analyzing…

stat.ME2023

Should I Stop or Should I Go: Early Stopping with Heterogeneous Populations

Hammaad Adam, Fan Yin, Huibin +5

Randomized experiments often need to be stopped prematurely due to the treatment having an unintended harmful effect. Existing methods that determine when to stop an experiment ear…

cs.HC2026

Fairness-in-the-Workflow: How Machine Learning Practitioners at Big Tech Companies Approach Fairness in Recommender Systems

Jing Nathan Yan, Emma Harvey, Junxiong Wang +2

Recommender systems (RS), which are widely deployed across high-stakes domains, are susceptible to biases that can cause large-scale societal impacts. Researchers have proposed met…