activity
20192026
most citedBeyond Preferences in AI Alignment

31 citations · 95 across the 22 of their papers we have counts for

collaborators
Showing 2024Show all

7 papers · 1 filter

cs.AI2024★ 31 cited

Beyond Preferences in AI Alignment

Tan Zhi-Xuan, Micah Carroll, Matija Franklin +1

The dominant practice of AI alignment assumes (1) that preferences are an adequate representation of human values, (2) that human rationality can be understood in terms of maximizi…

cs.CL2024

Understanding Epistemic Language with a Language-augmented Bayesian Theory of Mind

Lance Ying, Tan Zhi-Xuan, Lionel Wong +2

How do people understand and evaluate claims about others' beliefs, even though these beliefs cannot be directly observed? In this paper, we introduce a cognitive model of epistemi…

cs.AI2024

Infinite Ends from Finite Samples: Open-Ended Goal Inference as Top-Down Bayesian Filtering of Bottom-Up Proposals

Tan Zhi-Xuan, Gloria Kang, Vikash Mansinghka +1

The space of human goals is tremendously vast; and yet, from just a few moments of watching a scene or reading a story, we seem to spontaneously infer a range of plausible motivati…

cs.HC2024★ 2 cited

Building Machines that Learn and Think with People

Katherine M. Collins, Ilia Sucholutsky, Umang Bhatt +10

What do we want from machine intelligence? We envision machines that are not just tools for thought, but partners in thought: reasonable, insightful, knowledgeable, reliable, and t…

cs.AI2024★ 12 cited

Towards Guaranteed Safe AI: A Framework for Ensuring Robust and Reliable AI Systems

David "davidad" Dalrymple, Joar Skalse, Yoshua Bengio +14

Ensuring that AI systems reliably and robustly avoid harmful or dangerous behaviours is a crucial challenge, especially for AI systems with a high degree of autonomy and general in…

cs.AI2024★ 3 cited

Pragmatic Instruction Following and Goal Assistance via Cooperative Language-Guided Inverse Planning

Tan Zhi-Xuan, Lance Ying, Vikash Mansinghka +1

People often give instructions whose meaning is ambiguous without further context, expecting that their actions or goals will disambiguate their intentions. How can we build assist…