7 papers
Using AI Agents to Automate Black-Box Audits of Personalization Algorithms at Scale
Alessandro Morosini, Sarah H. Cen, Andrew Ilyas +3
Personalization algorithms determine what content users encounter on online platforms. Auditing these systems is difficult because independent auditors have only black-box access t…
DataMIL: Selecting Data for Robot Imitation Learning with Datamodels
Shivin Dass, Alaa Khaddaj, Logan Engstrom +3
Recently, the robotics community has amassed ever larger and more diverse datasets to train generalist policies. However, while these policies achieve strong mean performance acros…
Large-Scale, Longitudinal Study of Large Language Models During the 2024 US Election Season
Sarah H. Cen, Andrew Ilyas, Hedi Driss +4
The 2024 US presidential election is the first major contest to occur in the US since the popularization of large language models (LLMs). Building on lessons from earlier shifts in…
AI Supply Chains: An Emerging Ecosystem of AI Actors, Products, and Services
Aspen Hopkins, Sarah H. Cen, Andrew Ilyas +3
The widespread adoption of AI in recent years has led to the emergence of AI supply chains: complex networks of AI actors contributing models, datasets, and more to the development…
Optimizing ML Training with Metagradient Descent
Logan Engstrom, Andrew Ilyas, Benjamin Chen +3
A major challenge in training large-scale machine learning models is configuring the training process to maximize model performance, i.e., finding the best training setup from a va…
Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
Bowen Baker, Joost Huizinga, Leo Gao +6
Mitigating reward hacking--where AI systems misbehave due to flaws or misspecifications in their learning objectives--remains a key challenge in constructing capable and aligned mo…