9 citations · 11 across the 3 of their papers we have counts for
5 papers
Language and Experience: A Computational Model of Social Learning in Complex Tasks
Cédric Colas, Tracey Mills, Ben Prystawski +4
The ability to combine linguistic guidance from others with direct experience is central to human development, enabling safe and rapid learning in new environments. How do people i…
Bayesian Preference Elicitation with Language Models
Kunal Handa, Yarin Gal, Ellie Pavlick +4
Aligning AI systems to users' interests requires understanding and incorporating humans' complex values and preferences. Recently, language models (LMs) have been used to gather in…
Codebook Features: Sparse and Discrete Interpretability for Neural Networks
Alex Tamkin, Mohammad Taufeeque, Noah D. Goodman
Understanding neural networks is challenging in part because of the dense, continuous nature of their hidden states. We explore whether we can train neural networks to have hidden…
Eliciting Human Preferences with Language Models
Belinda Z. Li, Alex Tamkin, Noah Goodman +1
Language models (LMs) can be directed to perform target tasks by using labeled examples or natural language prompts. But selecting examples or writing prompts for can be challengin…
Social Contract AI: Aligning AI Assistants with Implicit Group Norms
Jan-Philipp Fränken, Sam Kwok, Peixuan Ye +6
We explore the idea of aligning an AI assistant by inverting a model of users' (unknown) preferences from observed interactions. To validate our proposal, we run proof-of-concept s…