7 papers
Healthcare LLM Benchmarks Are Only as Good as Their Explicit Assumptions
Naveen Raman, Santiago Cortes-Gomez, Mateo Dulce Rubio +2
Benchmarks are necessary for healthcare evaluation, but are not sufficient for predicting deployment performance. Our position is that the evaluation--deployment gap arises not bec…
The Limits of AI-Driven Allocation: Optimal Screening under Aleatoric Uncertainty
Santiago Cortes-Gomez, Mateo Dulce Rubio, Carlos Patino +1
The rise of machine learning has shifted targeted resource allocation in policy and humanitarian settings toward algorithmic targeting based on predicted risk scores. This approach…
When Agents Say One Thing and Do Another: Validating Elicited Beliefs from LLMs
Khurram Yamin, Jingjing Tang, Santiago Cortes-Gomez +3
Large language models (LLMs) are increasingly deployed in high-stakes settings where good decisions require forming beliefs over the probability of unknown outcomes. However, it is…
Predicting Language Models' Success at Zero-Shot Probabilistic Prediction
Kevin Ren, Santiago Cortes-Gomez, Carlos Miguel Patiño +7
Recent work has investigated the capabilities of large language models (LLMs) as zero-shot models for generating individual-level characteristics (e.g., to serve as risk models or…
Auditing Fairness by Betting
Ben Chugg, Santiago Cortes-Gomez, Bryan Wilder +1
We provide practical, efficient, and nonparametric methods for auditing the fairness of deployed classification and regression models. Whereas previous work relies on a fixed-sampl…
Utility-Directed Conformal Prediction: A Decision-Aware Framework for Actionable Uncertainty Quantification
Santiago Cortes-Gomez, Carlos Patiño, Yewon Byun +3
Interest has been growing in decision-focused machine learning methods which train models to account for how their predictions are used in downstream optimization problems. Doing s…