5 papers
Healthcare LLM Benchmarks Are Only as Good as Their Explicit Assumptions
Naveen Raman, Santiago Cortes-Gomez, Mateo Dulce Rubio +2
Benchmarks are necessary for healthcare evaluation, but are not sufficient for predicting deployment performance. Our position is that the evaluation--deployment gap arises not bec…
The Limits of AI-Driven Allocation: Optimal Screening under Aleatoric Uncertainty
Santiago Cortes-Gomez, Mateo Dulce Rubio, Carlos Patino +1
The rise of machine learning has shifted targeted resource allocation in policy and humanitarian settings toward algorithmic targeting based on predicted risk scores. This approach…
Sequentially Auditing Differential Privacy
Tomás González, Mateo Dulce-Rubio, Aaditya Ramdas +1
We propose a practical sequential test for auditing differential privacy guarantees of black-box mechanisms. The test processes streams of mechanisms' outputs providing anytime-val…
SESGO: Spanish Evaluation of Stereotypical Generative Outputs
Melissa Robles, Catalina Bernal, Denniss Raigoso +1
This paper addresses the critical gap in evaluating bias in multilingual Large Language Models (LLMs), with a specific focus on Spanish language within culturally-aware Latin Ameri…
Conformal Mixed-Integer Constraint Learning with Feasibility Guarantees
Daniel Ovalle, Lorenz T. Biegler, Ignacio E. Grossmann +2
We propose Conformal Mixed-Integer Constraint Learning (C-MICL), a novel framework that provides probabilistic feasibility guarantees for data-driven constraints in optimization pr…