1 paper · 1 filter
Rauno Arike, Elizabeth Donoway, Henning Bartsch +1
As language models (LMs) are increasingly deployed as autonomous agents, their robust adherence to human-assigned objectives becomes crucial for safe operation. When these agents o…