5 papers
Position: We Need Practical AI Alignment Methods to Mirror Human Reasoning
Vijay Keswani, Breanna K. Nguyen, Cyrus Cousins +3
AI systems are increasingly employed as decision aids, decision delegates, or autonomous decision-makers. This position paper argues that in many settings, particularly high-stakes…
Towards Cognitively-Faithful Decision-Making Models to Improve AI Alignment
Cyrus Cousins, Vijay Keswani, Vincent Conitzer +3
Recent AI trends seek to align AI models to learned human-centric objectives, such as personal preferences, utility, or societal values. Using standard preference elicitation metho…
International AI Safety Report 2026
Yoshua Bengio, Stephen Clare, Carina Prunkl +89
The International AI Safety Report 2026 synthesises the current scientific evidence on the capabilities, emerging risks, and safety of general-purpose AI systems. The report series…
Moral Change or Noise? On Problems of Aligning AI With Temporally Unstable Human Feedback
Vijay Keswani, Cyrus Cousins, Breanna Nguyen +4
Alignment methods in moral domains seek to elicit moral preferences of human stakeholders and incorporate them into AI. This presupposes moral preferences as static targets, but su…
Can AI Model the Complexities of Human Moral Decision-Making? A Qualitative Study of Kidney Allocation Decisions
Vijay Keswani, Vincent Conitzer, Walter Sinnott-Armstrong +3
A growing body of work in Ethical AI attempts to capture human moral judgments through simple computational models. The key question we address in this work is whether such simple…