3 papers
cs.LG2026
Redistribution-based Cost Inference Improves Sparse Safe Offline RL
Ebenezer Gelo, Geraud Nangue Tasse, Steven James +1
Safe offline RL typically assumes access to dense per-step cost annotations, but in practice supervisors provide only trajectory-level stop-feedback: a binary signal at the first u…
cs.AI2026
CORDA: A Benchmark for Hierarchical Harm-Centric Moral Reasoning in Large Language Models
Siddarth Singh, Victoria Williams, Simon Rosen +6
The key question in moral judgement is not simply whether someone chooses the "right" answer, but how they decide what matters most when moral principles conflict. Current evaluati…
cs.AI2026
MoralityGym: A Benchmark for Evaluating Hierarchical Moral Alignment in Sequential Decision-Making Agents
Simon Rosen, Siddarth Singh, Ebenezer Gelo +6
Evaluating moral alignment in agents navigating conflicting, hierarchically structured human norms is a critical challenge at the intersection of AI safety, moral philosophy, and c…