5 papers
Normative Robustness as a Frontier for Non-Verifiable Reasoning in LLMs
Elizaveta Tennant, Benjamin Henke, Anita Keshmirian +5
As LLMs increasingly serve in advisory and deliberative roles, users rely on them for non-verifiable reasoning in domains lacking objective ground truths. However, traditional eval…
Dynamics of Moral Behavior in Heterogeneous Populations of Learning Agents
Elizaveta Tennant, Stephen Hailes, Mirco Musolesi
Growing concerns about safety and alignment of AI systems highlight the importance of embedding moral capabilities in artificial agents: a promising solution is the use of learning…
Evaluating Generalization Capabilities of LLM-Based Agents in Mixed-Motive Scenarios Using Concordia
Chandler Smith, Marwa Abdulhai, Manfred Diaz +83
Large Language Model (LLM) agents have demonstrated impressive capabilities for social interaction and are increasingly being deployed in situations where they might engage with bo…
Moral Alignment for LLM Agents
Elizaveta Tennant, Stephen Hailes, Mirco Musolesi
Decision-making agents based on pre-trained Large Language Models (LLMs) are increasingly being deployed across various domains of human activity. While their applications are curr…
Hybrid Approaches for Moral Value Alignment in AI Agents: a Manifesto
Elizaveta Tennant, Stephen Hailes, Mirco Musolesi
Increasing interest in ensuring the safety of next-generation Artificial Intelligence (AI) systems calls for novel approaches to embedding morality into autonomous agents. This goa…