papers

Publications (47)

cs.IR2023

Reconciling the accuracy-diversity trade-off in recommendations

Kenny Peng, Manish Raghavan, Emma Pierson +2

In recommendation settings, there is an apparent trade-off between the goals of accuracy (to recommend items a user is most likely to want) and diversity (to recommend items repres…

cs.AI2024

Shaping AI's Impact on Billions of Lives

Mariano-Florentino Cuéllar, Jeff Dean, Finale Doshi-Velez +6

Artificial Intelligence (AI), like any transformative technology, has the potential to be a double-edged sword, leading either toward significant advancements or detrimental outcom…

cs.CL2024

Annotation alignment: Comparing LLM and human annotations of conversational safety

Rajiv Movva, Pang Wei Koh, Emma Pierson

Do LLMs align with human perceptions of safety? We study this question via annotation alignment, the extent to which LLMs and humans agree when annotating the safety of user-chatbo…

stat.AP2024

Testing for racial bias using inconsistent perceptions of race

Nora Gera, Emma Pierson

Tests for racial bias commonly assess whether two people of different races are treated differently. A fundamental challenge is that, because two people may differ in many ways, fa…

cs.CY2020

Ethical Machine Learning in Health Care

Irene Y. Chen, Emma Pierson, Sherri Rose +3

The use of machine learning (ML) in health care raises numerous ethical concerns, especially as models can amplify existing health inequities. Here, we outline ethical consideratio…

cs.LG2022

How do Authors' Perceptions of their Papers Compare with Co-authors' Perceptions and Peer-review Decisions?

Charvi Rastogi, Ivan Stelmakh, Alina Beygelzimer +7

How do author perceptions match up to the outcomes of the peer-review process and perceptions of others? In a top-tier computer science conference (NeurIPS 2021) with more than 23,…

cs.LG2024

Choosing the Right Weights: Balancing Value, Strategy, and Noise in Recommender Systems

Smitha Milli, Emma Pierson, Nikhil Garg

Many recommender systems optimize a linear weighting of different user behaviors, such as clicks, likes, and shares. We analyze the optimal choice of weights from the perspectives…

q-bio.GN2018

SIMLR: A Tool for Large-Scale Genomic Analyses by Multi-Kernel Learning

Bo Wang, Daniele Ramazzotti, Luca De Sano +3

We here present SIMLR (Single-cell Interpretation via Multi-kernel LeaRning), an open-source tool that implements a novel framework to learn a sample-to-sample similarity measure f…

cs.CY2026

Inferring fine-grained migration patterns across the United States

Gabriel Agostini, Rachel Young, Maria Fitzpatrick +2

Fine-grained migration data illuminate demographic, environmental, and health phenomena. However, United States migration data have serious drawbacks: public data lack spatial gran…

cs.LG2026

Position: Use Sparse Autoencoders to Discover Unknowns

Kenny Peng, Rajiv Movva, Jon Kleinberg +2

While sparse autoencoders (SAEs) have generated significant excitement, a series of negative results have added to skepticism about their usefulness. Here, we establish a conceptua…

cs.CL2026

What's In My Human Feedback? Learning Interpretable Descriptions of Preference Data

Rajiv Movva, Smitha Milli, Sewon Min +1

Human feedback can alter language models in unpredictable and undesirable ways, as practitioners lack a clear understanding of what feedback data encodes. While prior work studies…

cs.SI2018

Modeling Individual Cyclic Variation in Human Behavior

Emma Pierson, Tim Althoff, Jure Leskovec

Cycles are fundamental to human health and behavior. However, modeling cycles in time series data is challenging because in most cases the cycles are not labeled or directly observ…

cs.LG2024

Domain constraints improve risk prediction when outcome data is missing

Sidhika Balachandar, Nikhil Garg, Emma Pierson

Machine learning models are often trained to predict the outcome resulting from a human decision. For example, if a doctor decides to test a patient for disease, will the patient t…

cs.CY2026

In your own words: computationally identifying interpretable themes in free-text survey data

Jenny S Wang, Aliya Saperstein, Emma Pierson

Free-text survey responses can provide nuance often missed by structured questions, but remain difficult to statistically analyze. To address this, we introduce In Your Own Words,…

q-bio.QM2025

Disentangling Proxies of Demographic Adjustments in Clinical Equations

Aashna P. Shah, James A. Diao, Emma Pierson +2

The use of coarse demographic adjustments in clinical equations has been increasingly scrutinized. In particular, adjustments for race have sparked significant debate with several…

cs.LG2025

Learning Disease Progression Models That Capture Health Disparities

Erica Chiang, Divya Shanmugam, Ashley N. Beecy +4

Disease progression models are widely used to inform the diagnosis and treatment of many progressive diseases. However, a significant limitation of existing models is that they do…

cs.HC2022

Trucks Don't Mean Trump: Diagnosing Human Error in Image Analysis

J. D. Zamfirescu-Pereira, Jerry Chen, Emily Wen +3

Algorithms provide powerful tools for detecting and dissecting human bias and error. Here, we develop machine learning methods to to analyze how humans err in a particular high-sta…

cs.LG2025

Bayesian Modeling of Zero-Shot Classifications for Urban Flood Detection

Matt Franchi, Nikhil Garg, Wendy Ju +1

Street scene datasets, collected from Street View or dashboard cameras, offer a promising means of detecting urban objects and incidents like street flooding. However, a major chal…

cs.LG2026

Strategic Feature Selection

Jivat Neet Kaur, Pratik Patil, Divya Shanmugam +6

When algorithmic predictors inform resource allocation in high-stakes domains such as healthcare, these predictors must account for strategic manipulation of input features. The ty…

cs.CY2025

Using large language models to promote health equity

Emma Pierson, Divya Shanmugam, Rajiv Movva +12

Advances in large language models (LLMs) have driven an explosion of interest about their societal impacts. Much of the discourse around how they will impact social equity has been…

cs.LG2020

Concept Bottleneck Models

Pang Wei Koh, Thao Nguyen, Yew Siang Tang +4

We seek to learn models that we can interact with using high-level concepts: if the model did not think there was a bone spur in the x-ray, would it still predict severe arthritis?…

cs.DL2024

Topics, Authors, and Institutions in Large Language Model Research: Trends from 17K arXiv Papers

Rajiv Movva, Sidhika Balachandar, Kenny Peng +3

Large language models (LLMs) are dramatically influencing AI research, spurring discussions on what has changed so far and how to shape the field's future. To clarify such question…

cs.CL2025

Sparse Autoencoders for Hypothesis Generation

Rajiv Movva, Kenny Peng, Nikhil Garg +2

We describe HypotheSAEs, a general method to hypothesize interpretable relationships between text data (e.g., headlines) and a target variable (e.g., clicks). HypotheSAEs has three…

stat.AP2020

Assessing racial inequality in COVID-19 testing with Bayesian threshold tests

Emma Pierson

There are racial disparities in the COVID-19 test positivity rate, suggesting that minorities may be under-tested. Here, drawing on the literature on statistically assessing racial…

stat.ML2018

Fast Threshold Tests for Detecting Discrimination

Emma Pierson, Sam Corbett-Davies, Sharad Goel

Threshold tests have recently been proposed as a useful method for detecting bias in lending, hiring, and policing decisions. For example, in the case of credit extensions, these t…

cs.CY2023

Coarse race data conceals disparities in clinical risk score performance

Rajiv Movva, Divya Shanmugam, Kaihua Hou +4

Healthcare data in the United States often records only a patient's coarse race group: for example, both Indian and Chinese patients are typically coded as "Asian." It is unknown,…

cs.CL2024

MediQ: Question-Asking LLMs and a Benchmark for Reliable Interactive Clinical Reasoning

Shuyue Stella Li, Vidhisha Balachandran, Shangbin Feng +4

Users typically engage with LLMs interactively, yet most existing benchmarks evaluate them in a static, single-turn format, posing reliability concerns in interactive scenarios. We…

cs.LG2024

Generative AI in Medicine

Divya Shanmugam, Monica Agrawal, Rajiv Movva +4

The increased capabilities of generative AI have dramatically expanded its possible use cases in medicine. We provide a comprehensive overview of generative AI use cases for clinic…

cs.LG2025

Evaluating multiple models using labeled and unlabeled data

Divya Shanmugam, Shuvom Sadhuka, Manish Raghavan +3

It remains difficult to evaluate machine learning classifiers in the absence of a large, labeled dataset. While labeled data can be prohibitively expensive or impossible to obtain,…

cs.CY2025

Advancing Science- and Evidence-based AI Policy

Rishi Bommasani, Sanjeev Arora, Jennifer Chayes +17

AI policy should advance AI innovation by ensuring that its potential benefits are responsibly realized and widely shared. To achieve this, AI policymaking should place a premium o…

cs.CY2018

Demographics and discussion influence views on algorithmic fairness

Emma Pierson

The field of algorithmic fairness has highlighted ethical questions which may not have purely technical answers. For example, different algorithmic fairness constraints are often i…

cs.CY2025

LLMs generate structurally realistic social networks but overestimate political homophily

Serina Chang, Alicja Chaszczewicz, Emma Wang +3

Generating social networks is essential for many applications, such as epidemic modeling and social simulations. The emergence of generative AI, especially large language models (L…

stat.AP2019

Predicting pregnancy using large-scale data from a women's health tracking mobile application

Bo Liu, Shuyang Shi, Yongshang Wu +4

Predicting pregnancy has been a fundamental problem in women's health for more than 50 years. Previous datasets have been collected via carefully curated medical studies, but the r…

stat.AP2017

A large-scale analysis of racial disparities in police stops across the United States

Emma Pierson, Camelia Simoiu, Jan Overgoor +4

To assess racial disparities in police interactions with the public, we compiled and analyzed a dataset detailing over 60 million state patrol stops conducted in 20 U.S. states bet…

cs.LG2025

A Bayesian Model for Multi-stage Censoring

Shuvom Sadhuka, Sophia Lin, Bonnie Berger +1

Many sequential decision settings in healthcare feature funnel structures characterized by a series of stages, such as screenings or evaluations, where the number of patients who a…

cs.LG2025

Urban Incident Prediction with Graph Neural Networks: Integrating Government Ratings and Crowdsourced Reports

Sidhika Balachandar, Shuvom Sadhuka, Bonnie Berger +2

Graph neural networks (GNNs) are widely used in urban spatiotemporal forecasting, such as predicting infrastructure problems. In this setting, government officials wish to know in…

cs.CY2023

A Bayesian Spatial Model to Correct Under-Reporting in Urban Crowdsourcing

Gabriel Agostini, Emma Pierson, Nikhil Garg

Decision-makers often observe the occurrence of events through a reporting process. City governments, for example, rely on resident reports to find and then resolve urban infrastru…

cs.CY2023

Detecting disparities in police deployments using dashcam data

Matt Franchi, J. D. Zamfirescu-Pereira, Wendy Ju +1

Large-scale policing data is vital for detecting inequity in police behavior and policing algorithms. However, one important type of policing data remains largely unavailable withi…

cs.CY2026

Three Years of r/ChatGPT: Societal Impact Evaluations from Social Media Data

Jessica Dai, Sean Garcia, Emma Pierson +2

ChatGPT was launched on November 30, 2022; the r/ChatGPT subreddit was created just one day later. Since then, chatbot-based AI products have gone from niche proofs-of-concept to w…

cs.LG2021

WILDS: A Benchmark of in-the-Wild Distribution Shifts

Pang Wei Koh, Shiori Sagawa, Henrik Marklund +20

Distribution shifts -- where the training distribution differs from the test distribution -- can substantially degrade the accuracy of machine learning (ML) systems deployed in the…

cs.LG2024

Recent Advances, Applications, and Open Challenges in Machine Learning for Health: Reflections from Research Roundtables at ML4H 2023 Symposium

Hyewon Jeong, Sarah Jabbour, Yuzhe Yang +40

The third ML4H symposium was held in person on December 10, 2023, in New Orleans, Louisiana, USA. The symposium included research roundtable sessions to foster discussions between…

cs.SI2023

Human mobility networks reveal increased segregation in large cities

Hamed Nilforoshan, Wenli Looi, Emma Pierson +7

A long-standing expectation is that large, dense, and cosmopolitan areas support socioeconomic mixing and exposure between diverse individuals. It has been difficult to assess this…

cs.CY2023

Quantifying disparities in intimate partner violence: a machine learning method to correct for underreporting

Divya Shanmugam, Kaihua Hou, Emma Pierson

Estimating the prevalence of a medical condition, or the proportion of the population in which it occurs, is a fundamental problem in healthcare and public health. Accurate estimat…

cs.LG2019

Inferring Multidimensional Rates of Aging from Cross-Sectional Data

Emma Pierson, Pang Wei Koh, Tatsunori Hashimoto +4

Modeling how individuals evolve over time is a fundamental problem in the natural and social sciences. However, existing datasets are often cross-sectional with each individual obs…

cs.CY2024

Participation in the age of foundation models

Harini Suresh, Emily Tseng, Meg Young +3

Growing interest and investment in the capabilities of foundation models has positioned such systems to impact a wide array of public services. Alongside these opportunities is the…

cs.CY2017

Algorithmic decision making and the cost of fairness

Sam Corbett-Davies, Emma Pierson, Avi Feller +2

Algorithms are now regularly used to decide whether defendants awaiting trial are too dangerous to be released back into the community. In some cases, black defendants are substant…

cs.LG2025

Reflections from Research Roundtables at the Conference on Health, Inference, and Learning (CHIL) 2025

Emily Alsentzer, Marie-Laure Charpignon, Bill Chen +90

The 6th Annual Conference on Health, Inference, and Learning (CHIL 2025), hosted by the Association for Health Learning and Inference (AHLI), was held in person on June 25-27, 2025…