193 citations · 752 across the 18 of their papers we have counts for
3 papers · 1 filter
Fine-tuning language models to find agreement among humans with diverse preferences
Michiel A. Bakker, Martin J. Chadwick, Hannah R. Sheahan +8
Recent work in large language modeling (LLMs) has used fine-tuning to align outputs with the preferences of a prototypical user. This work assumes that human preferences are static…
The Frost Hollow Experiments: Pavlovian Signalling as a Path to Coordination and Communication Between Agents
Patrick M. Pilarski, Andrew Butcher, Elnaz Davoodi +7
Learned communication between agents is a powerful tool when approaching decision-making problems that are hard to overcome by any single agent in isolation. However, continual coo…
HCMD-zero: Learning Value Aligned Mechanisms from Data
Jan Balaguer, Raphael Koster, Ari Weinstein +4
Artificial learning agents are mediating a larger and larger number of interactions among humans, firms, and organizations, and the intersection between mechanism design and machin…