1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.AI2026
MoralityGym: A Benchmark for Evaluating Hierarchical Moral Alignment in Sequential Decision-Making Agents
Simon Rosen, Siddarth Singh, Ebenezer Gelo +6
Evaluating moral alignment in agents navigating conflicting, hierarchically structured human norms is a critical challenge at the intersection of AI safety, moral philosophy, and c…
cs.AI2023★ 1 cited
MinePlanner: A Benchmark for Long-Horizon Planning in Large Minecraft Worlds
William Hill, Ireton Liu, Anita De Mello Koch +4
We propose a new benchmark for planning tasks based on the Minecraft game. Our benchmark contains 45 tasks overall, but also provides support for creating both propositional and nu…