238 citations · 242 across the 3 of their papers we have counts for
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2024★ 3 cited
Holistic Safety and Responsibility Evaluations of Advanced AI Models
Laura Weidinger, Joslyn Barnhart, Jenny Brennan +16
Safety and responsibility evaluations of advanced AI models are a critical but developing field of research and practice. In the development of Google DeepMind's advanced AI models…
cs.AI2023★ 1 cited
Doing the right thing for the right reason: Evaluating artificial moral cognition by probing cost insensitivity
Yiran Mao, Madeline G. Reinecke, Markus Kunesch +4
Is it possible to evaluate the moral cognition of complex artificial agents? In this work, we take a look at one aspect of morality: `doing the right thing for the right reasons.'…