3 papers
cs.AI2026
A Descriptive and Normative Theory of Human Beliefs in RLHF
Sylee Dandekar, Shripad Deshmukh, Frank Chiu +2
Human preferences in RLHF are typically modeled as a function of the human's reward function or corresponding optimal state-action values. In this work, we propose that human belie…
cs.LG2025
Full-Stack Alignment: Co-Aligning AI and Institutions with Thick Models of Value
Joe Edelman, Tan Zhi-Xuan, Ryan Lowe +30
Beneficial societal outcomes cannot be guaranteed by aligning individual AI systems with the intentions of their operators or users. Even an AI system that is perfectly aligned to…
cs.CL2025
Modeling Future Conversation Turns to Teach LLMs to Ask Clarifying Questions
Michael J. Q. Zhang, W. Bradley Knox, Eunsol Choi
Large language models (LLMs) must often respond to highly ambiguous user requests. In such cases, the LLM's best response may be to ask a clarifying question to elicit more informa…