28 citations · 31 across the 4 of their papers we have counts for
4 papers
Uncovering Latent Human Wellbeing in Language Model Embeddings
Pedro Freire, ChengCheng Tan, Adam Gleave +2
Do language models implicitly learn a concept of human wellbeing? We explore this through the ETHICS Utilitarianism task, assessing if scaling enhances pretrained models' represent…
Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark
Alexander Pan, Jun Shern Chan, Andy Zou +7
Artificial agents have traditionally been trained to maximize reward, which may incentivize power-seeking and deception, analogous to how next-token prediction in language models (…
For Learning in Symmetric Teams, Local Optima are Global Nash Equilibria
Scott Emmons, Caspar Oesterheld, Andrew Critch +2
Although it has been known since the 1970s that a globally optimal strategy profile in a common-payoff game is a Nash equilibrium, global optimality is a strict requirement that li…
RvS: What is Essential for Offline RL via Supervised Learning?
Scott Emmons, Benjamin Eysenbach, Ilya Kostrikov +1
Recent work has shown that supervised learning alone, without temporal difference (TD) learning, can be remarkably effective for offline RL. When does this hold true, and which alg…