3 citations · 7 across the 5 of their papers we have counts for
Showing cs.LGShow all
3 papers · 1 filter
cs.LG2025
Full-Stack Alignment: Co-Aligning AI and Institutions with Thick Models of Value
Joe Edelman, Tan Zhi-Xuan, Ryan Lowe +30
Beneficial societal outcomes cannot be guaranteed by aligning individual AI systems with the intentions of their operators or users. Even an AI system that is perfectly aligned to…
cs.LG2024
Sabotage Evaluations for Frontier Models
Joe Benton, Misha Wagner, Eric Christiansen +13
Sufficiently capable models could subvert human oversight and decision-making in important contexts. For example, in the context of AI development, models could covertly sabotage e…
cs.LG2024
Experts Don't Cheat: Learning What You Don't Know By Predicting Pairs
Daniel D. Johnson, Daniel Tarlow, David Duvenaud +1
Identifying how much a model knows about the stochastic real-world process it was trained on is important to ensure it avoids producing incorrect o…