28 citations · 50 across the 3 of their papers we have counts for
3 papers
cs.CY2024
Responsible Reporting for Frontier AI Development
Noam Kolt, Markus Anderljung, Joslyn Barnhart +7
Mitigating the risks from frontier AI systems requires up-to-date and reliable information about those systems. Organizations that develop and deploy frontier systems have signific…
cs.LG2023★ 28 cited
Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark
Alexander Pan, Jun Shern Chan, Andy Zou +7
Artificial agents have traditionally been trained to maximize reward, which may incentivize power-seeking and deception, analogous to how next-token prediction in language models (…
cs.CY2023★ 22 cited
Artificial Influence: An Analysis Of AI-Driven Persuasion
Matthew Burtell, Thomas Woodside
Persuasion is a key aspect of what it means to be human, and is central to business, politics, and other endeavors. Advancements in artificial intelligence (AI) have produced AI sy…