9 citations · 12 across the 5 of their papers we have counts for
Showing 2023 · cs.CLShow all
2 papers · 2 filters
cs.CL2023★ 9 cited
Eliciting Human Preferences with Language Models
Belinda Z. Li, Alex Tamkin, Noah Goodman +1
Language models (LMs) can be directed to perform target tasks by using labeled examples or natural language prompts. But selecting examples or writing prompts for can be challengin…
cs.CL2023★ 1 cited
Social Contract AI: Aligning AI Assistants with Implicit Group Norms
Jan-Philipp Fränken, Sam Kwok, Peixuan Ye +6
We explore the idea of aligning an AI assistant by inverting a model of users' (unknown) preferences from observed interactions. To validate our proposal, we run proof-of-concept s…