28 citations · 28 across the 1 of their papers we have counts for
1 paper
Alexander Pan, Jun Shern Chan, Andy Zou +7
Artificial agents have traditionally been trained to maximize reward, which may incentivize power-seeking and deception, analogous to how next-token prediction in language models (…