Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Full-Stack Alignment: Co-Aligning AI and Institutions with Thick Models of Value
Joe Edelman, Tan Zhi-Xuan, Ryan Lowe +30
Beneficial societal outcomes cannot be guaranteed by aligning individual AI systems with the intentions of their operators or users. Even an AI system that is perfectly aligned to…
cs.LG2024
Social Choice Should Guide AI Alignment in Dealing with Diverse Human Feedback
Vincent Conitzer, Rachel Freedman, Jobst Heitzig +9
Foundation models such as GPT-4 are fine-tuned to avoid unsafe or otherwise problematic behavior, such as helping to commit crimes or producing racist text. One approach to fine-tu…