71 citations · 71 across the 1 of their papers we have counts for
2 papers
cs.AI2023★ 1 cited
Doing the right thing for the right reason: Evaluating artificial moral cognition by probing cost insensitivity
Yiran Mao, Madeline G. Reinecke, Markus Kunesch +4
Is it possible to evaluate the moral cognition of complex artificial agents? In this work, we take a look at one aspect of morality: `doing the right thing for the right reasons.'…
cs.CL2021★ 71 cited
Ethical and social risks of harm from Language Models
Laura Weidinger, John Mellor, Maribeth Rauh +20
This paper aims to help structure the risk landscape associated with large-scale Language Models (LMs). In order to foster advances in responsible innovation, an in-depth understan…