9 papers
Using AI Agents to Automate Black-Box Audits of Personalization Algorithms at Scale
Alessandro Morosini, Sarah H. Cen, Andrew Ilyas +3
Personalization algorithms determine what content users encounter on online platforms. Auditing these systems is difficult because independent auditors have only black-box access t…
DataMIL: Selecting Data for Robot Imitation Learning with Datamodels
Shivin Dass, Alaa Khaddaj, Logan Engstrom +3
Recently, the robotics community has amassed ever larger and more diverse datasets to train generalist policies. However, while these policies achieve strong mean performance acros…
Justified or Just Convincing? Error Verifiability as a Dimension of LLM Quality
Xiaoyuan Zhu, Kimberly Le Truong, Riccardo Fogliato +8
As LLMs are deployed in high-stakes settings, users must judge the correctness of individual responses, often relying on model-generated justifications such as reasoning chains or…
Probably Approximately Correct Labels
Emmanuel J. Candès, Andrew Ilyas, Tijana Zrnic
Obtaining high-quality labeled datasets is often costly, requiring either human annotation or expensive experiments. In theory, powerful pre-trained AI models provide an opportunit…
Large-Scale, Longitudinal Study of Large Language Models During the 2024 US Election Season
Sarah H. Cen, Andrew Ilyas, Hedi Driss +4
The 2024 US presidential election is the first major contest to occur in the US since the popularization of large language models (LLMs). Building on lessons from earlier shifts in…
AI Supply Chains: An Emerging Ecosystem of AI Actors, Products, and Services
Aspen Hopkins, Sarah H. Cen, Andrew Ilyas +3
The widespread adoption of AI in recent years has led to the emergence of AI supply chains: complex networks of AI actors contributing models, datasets, and more to the development…