4 papers
Hedging and Non-Affirmation: Quantifying LLM Alignment on Questions of Human Rights
Rafiya Javed, Cassandra Parent, Jackie Kay +9
Hedging and non-affirmation are behaviors exhibited by large language models (LLMs) that limit the clear endorsement of specific statements. While these behaviors are desirable in…
Value Profiles for Encoding Human Variation
Taylor Sorensen, Pushkar Mishra, Roma Patel +6
Modelling human variation in rating tasks is crucial for personalization, pluralistic model alignment, and computational social science. We propose representing individuals using n…
Virtual Agent Economies
Nenad Tomasev, Matija Franklin, Joel Z. Leibo +4
The rapid adoption of autonomous AI agents is giving rise to a new economic layer where agents transact and coordinate at scales and speeds beyond direct human oversight. We propos…
Multi-Agent Risks from Advanced AI
Lewis Hammond, Alan Chan, Jesse Clifton +41
The rapid development of advanced AI agents and the imminent deployment of many instances of these agents will give rise to multi-agent systems of unprecedented complexity. These s…