31 citations · 95 across the 22 of their papers we have counts for
7 papers · 1 filter
Beyond Preferences in AI Alignment
Tan Zhi-Xuan, Micah Carroll, Matija Franklin +1
The dominant practice of AI alignment assumes (1) that preferences are an adequate representation of human values, (2) that human rationality can be understood in terms of maximizi…
Understanding Epistemic Language with a Language-augmented Bayesian Theory of Mind
Lance Ying, Tan Zhi-Xuan, Lionel Wong +2
How do people understand and evaluate claims about others' beliefs, even though these beliefs cannot be directly observed? In this paper, we introduce a cognitive model of epistemi…
Infinite Ends from Finite Samples: Open-Ended Goal Inference as Top-Down Bayesian Filtering of Bottom-Up Proposals
Tan Zhi-Xuan, Gloria Kang, Vikash Mansinghka +1
The space of human goals is tremendously vast; and yet, from just a few moments of watching a scene or reading a story, we seem to spontaneously infer a range of plausible motivati…
Building Machines that Learn and Think with People
Katherine M. Collins, Ilia Sucholutsky, Umang Bhatt +10
What do we want from machine intelligence? We envision machines that are not just tools for thought, but partners in thought: reasonable, insightful, knowledgeable, reliable, and t…
Towards Guaranteed Safe AI: A Framework for Ensuring Robust and Reliable AI Systems
David "davidad" Dalrymple, Joar Skalse, Yoshua Bengio +14
Ensuring that AI systems reliably and robustly avoid harmful or dangerous behaviours is a crucial challenge, especially for AI systems with a high degree of autonomy and general in…
Pragmatic Instruction Following and Goal Assistance via Cooperative Language-Guided Inverse Planning
Tan Zhi-Xuan, Lance Ying, Vikash Mansinghka +1
People often give instructions whose meaning is ambiguous without further context, expecting that their actions or goals will disambiguate their intentions. How can we build assist…