9 papers
Model Agreement via Anchoring
Eric Eaton, Surbhi Goel, Marcel Hussing +4
Numerous lines of aim to control -- the extent to which two machine learning models disagree in their predictions. We adopt a simple and standard noti…
Personalization Aids Pluralistic Alignment Under Competition
Natalie Collina, Surbhi Goel, Aaron Roth +1
Can competition among misaligned AI providers yield aligned outcomes for a diverse population of users, and what role does model personalization play? We study a setting where mult…
Emergent Alignment via Competition
Natalie Collina, Surbhi Goel, Aaron Roth +2
Aligning AI systems with human values remains a fundamental challenge, but does our inability to create perfectly aligned models preclude obtaining the benefits of alignment? We st…
Why Do Transformers Fail to Forecast Time Series In-Context?
Yufa Zhou, Yixiao Wang, Surbhi Goel +1
Time series forecasting (TSF) remains a challenging and largely unsolved problem in machine learning, despite significant recent efforts leveraging Large Language Models (LLMs), wh…
Probabilistic Stability Guarantees for Feature Attributions
Helen Jin, Anton Xue, Weiqiu You +2
Stability guarantees have emerged as a principled way to evaluate feature attributions, but existing certification methods rely on heavily smoothed classifiers and often produce co…
Conformal Language Model Reasoning with Coherent Factuality
Maxon Rubin-Toles, Maya Gambhir, Keshav Ramji +2
Language models are increasingly being used in important decision pipelines, so ensuring the correctness of their outputs is crucial. Recent work has proposed evaluating the "factu…