activity
20232025
most citedEliciting Human Preferences with Language Models

9 citations · 11 across the 3 of their papers we have counts for

collaborators

5 papers

cs.AI2025

Language and Experience: A Computational Model of Social Learning in Complex Tasks

Cédric Colas, Tracey Mills, Ben Prystawski +4

The ability to combine linguistic guidance from others with direct experience is central to human development, enabling safe and rapid learning in new environments. How do people i…

cs.CL2024

Bayesian Preference Elicitation with Language Models

Kunal Handa, Yarin Gal, Ellie Pavlick +4

Aligning AI systems to users' interests requires understanding and incorporating humans' complex values and preferences. Recently, language models (LMs) have been used to gather in…

cs.LG20232 cited

Codebook Features: Sparse and Discrete Interpretability for Neural Networks

Alex Tamkin, Mohammad Taufeeque, Noah D. Goodman

Understanding neural networks is challenging in part because of the dense, continuous nature of their hidden states. We explore whether we can train neural networks to have hidden…

cs.CL20239 cited

Eliciting Human Preferences with Language Models

Belinda Z. Li, Alex Tamkin, Noah Goodman +1

Language models (LMs) can be directed to perform target tasks by using labeled examples or natural language prompts. But selecting examples or writing prompts for can be challengin…

cs.CL2023

Social Contract AI: Aligning AI Assistants with Implicit Group Norms

Jan-Philipp Fränken, Sam Kwok, Peixuan Ye +6

We explore the idea of aligning an AI assistant by inverting a model of users' (unknown) preferences from observed interactions. To validate our proposal, we run proof-of-concept s…