most citedDo the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

28 citations · 29 across the 5 of their papers we have counts for

collaborators

5 papers

cs.AI2023

Improved Logical Reasoning of Language Models via Differentiable Symbolic Programming

Hanlin Zhang, Jiani Huang, Ziyang Li +2

Pre-trained large language models (LMs) struggle to perform logical reasoning reliably despite advances in scale and compositionality. In this work, we tackle this challenge throug…

cs.LG202328 cited

Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

Alexander Pan, Jun Shern Chan, Andy Zou +7

Artificial agents have traditionally been trained to maximize reward, which may incentivize power-seeking and deception, analogous to how next-token prediction in language models (…

cs.LG2023

Fuzziness-tuned: Improving the Transferability of Adversarial Examples

Xiangyuan Yang, Jie Lin, Hanlin Zhang +2

With the development of adversarial attacks, adversairal examples have been widely used to enhance the robustness of the training models on deep neural networks. Although considera…

cs.LG2022

Improving the Robustness and Generalization of Deep Neural Network with Confidence Threshold Reduction

Xiangyuan Yang, Jie Lin, Hanlin Zhang +2

Deep neural networks are easily attacked by imperceptible perturbation. Presently, adversarial training (AT) is the most effective method to enhance the robustness of the model aga…

cs.LG20221 cited

Stochastic Neural Networks with Infinite Width are Deterministic

Liu Ziyin, Hanlin Zhang, Xiangming Meng +3

This work theoretically studies stochastic neural networks, a main type of neural network in use. We prove that as the width of an optimized stochastic neural network tends to infi…