most citedLearning Representations that Enable Generalization in Assistive Tasks

4 citations · 9 across the 5 of their papers we have counts for

collaborators

13 papers

cs.LG20241 cited

Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2

Tom Lieberum, Senthooran Rajamanoharan, Arthur Conmy +7

Sparse autoencoders (SAEs) are an unsupervised method for learning a sparse decomposition of a neural network's latent representations into seemingly interpretable features. Despit…

cs.CR20241 cited

Adversaries Can Misuse Combinations of Safe Models

Erik Jones, Anca Dragan, Jacob Steinhardt

Developers try to evaluate whether an AI system can be misused by adversaries before releasing it; for example, they might test whether a model enables cyberoffense, user manipulat…

cs.LG202411 cited

Evaluating Frontier Models for Dangerous Capabilities

Mary Phuong, Matthew Aitchison, Elliot Catt +24

To understand the risks posed by a new AI system, we must understand what it can and cannot do. Building on prior work, we introduce a programme of new "dangerous capability" evalu…

cs.RO20241 cited

A Generalized Acquisition Function for Preference-based Reward Learning

Evan Ellis, Gaurav R. Ghosal, Stuart J. Russell +2

Preference-based reward learning is a popular technique for teaching robots and autonomous systems how a human user wants them to perform a task. Previous works have shown that act…

cs.LG20232 cited

Zero-Shot Goal-Directed Dialogue via RL on Imagined Conversations

Joey Hong, Sergey Levine, Anca Dragan

Large language models (LLMs) have emerged as powerful and general solutions to many natural language tasks. However, many of the most important applications of language generation…

cs.LG2023

Offline RL with Observation Histories: Analyzing and Improving Sample Complexity

Joey Hong, Anca Dragan, Sergey Levine

Offline reinforcement learning (RL) can in principle synthesize more optimal behavior from a dataset consisting only of suboptimal trials. One way that this can happen is by "stitc…