6 papers
CORDA: A Benchmark for Hierarchical Harm-Centric Moral Reasoning in Large Language Models
Siddarth Singh, Victoria Williams, Simon Rosen +6
The key question in moral judgement is not simply whether someone chooses the "right" answer, but how they decide what matters most when moral principles conflict. Current evaluati…
MoralityGym: A Benchmark for Evaluating Hierarchical Moral Alignment in Sequential Decision-Making Agents
Simon Rosen, Siddarth Singh, Ebenezer Gelo +6
Evaluating moral alignment in agents navigating conflicting, hierarchically structured human norms is a critical challenge at the intersection of AI safety, moral philosophy, and c…
Self-Supervised On-Policy Reinforcement Learning via Contrastive Proximal Policy Optimisation
Asim Osman, Sasha Abramowitz, Mark Bergh +13
Contrastive reinforcement learning (CRL) learns goal-conditioned Q-values through a contrastive objective over state-action and goal representations, removing the need for hand-cra…
Characterizing MARL for Energy Control: A Multi-KPI Benchmark on the CityLearn Environment
Aymen Khouja, Imen Jendoubi, Oumayma Mahjoub +4
The optimization of urban energy systems is crucial for the advancement of sustainable and resilient smart cities, which are becoming increasingly complex with multiple decision-ma…
Breaking the Performance Ceiling in Reinforcement Learning requires Inference Strategies
Felix Chalumeau, Daniel Rajaonarivonivelomanantsoa, Ruan de Kock +12
Reinforcement learning (RL) systems have countless applications, from energy-grid management to protein design. However, such real-world scenarios are often extremely difficult, co…
Oryx: a Scalable Sequence Model for Many-Agent Coordination in Offline MARL
Claude Formanek, Omayma Mahjoub, Louay Ben Nessir +10
A key challenge in offline multi-agent reinforcement learning (MARL) is achieving effective many-agent multi-step coordination in complex environments. In this work, we propose Ory…