5 papers
Latent Adversarial Regularization for Offline Preference Optimization
Enyi Jiang, Yibo Jacky Zhang, Yinglun Xu +3
Learning from human feedback typically relies on preference optimization that constrains policy updates through token-level regularization. However, preference optimization for lan…
Full-Stack Alignment: Co-Aligning AI and Institutions with Thick Models of Value
Joe Edelman, Tan Zhi-Xuan, Ryan Lowe +30
Beneficial societal outcomes cannot be guaranteed by aligning individual AI systems with the intentions of their operators or users. Even an AI system that is perfectly aligned to…
Preference Measurement Error, Concentration in Recommendation Systems, and Persuasion
Andreas Haupt
Algorithmic recommendation based on noisy preference measurement is prevalent in recommendation systems. This paper discusses the consequences of such recommendation on market conc…
Position: Machine Learning Conferences Should Establish a "Refutations and Critiques" Track
Rylan Schaeffer, Joshua Kazdan, Yegor Denisov-Blanch +11
Science progresses by iteratively advancing and correcting humanity's understanding of the world. In machine learning (ML) research, rapid advancements have led to an explosion of…
Scaling Human Judgment in Community Notes with LLMs
Haiwen Li, Soham De, Manon Revel +6
This paper argues for a new paradigm for Community Notes in the LLM era: an open ecosystem where both humans and LLMs can write notes, and the decision of which notes are helpful e…