activity
20232026
most citedGemma 2: Improving Open Language Models at a Practical Size

145 citations · 172 across the 23 of their papers we have counts for

collaborators
Showing cs.AIShow all

7 papers · 1 filter

cs.AI2025

Training LLM Agents to Empower Humans

Evan Ellis, Vivek Myers, Jens Tuyls +3

Assistive agents should not only take actions on behalf of a human, but also step out of the way and cede control when there are important decisions to be made. However, current me…

cs.AI2025

CTRL-Rec: Controlling Recommender Systems With Natural Language

Micah Carroll, Adeline Foote, Kevin Feng +4

When users are dissatisfied with recommendations from a recommender system, they often lack fine-grained controls for changing them. Large language models (LLMs) offer a solution b…

cs.AI20252 cited

Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety

Tomek Korbak, Mikita Balesni, Elizabeth Barnes +38

AI systems that "think" in human language offer a unique opportunity for AI safety: we can monitor their chains of thought (CoT) for the intent to misbehave. Like all other known A…

cs.AI20251 cited

An Approach to Technical AGI Safety and Security

Rohin Shah, Alex Irpan, Alexander Matt Turner +27

Artificial General Intelligence (AGI) promises transformative benefits but also presents significant risks. We develop an approach to address the risk of harms consequential enough…

cs.AI2025

AssistanceZero: Scalably Solving Assistance Games

Cassidy Laidlaw, Eli Bronstein, Timothy Guo +5

Assistance games are a promising alternative to reinforcement learning from human feedback (RLHF) for training AI assistants. Assistance games resolve key drawbacks of RLHF, such a…

cs.AI2024

Learning to Assist Humans without Inferring Rewards

Vivek Myers, Evan Ellis, Sergey Levine +2

Assistive agents should make humans' lives easier. Classically, such assistance is studied through the lens of inverse reinforcement learning, where an assistive agent (e.g., a cha…