activity
20242026
collaborators

6 papers

cs.GT2026

Computationally Efficient Collaborative Communication Via Regularity-Based Coarsening

Mark Bedaywi, Scott Emmons, Nika Haghtalab +1

Our results show that the existence of a short high-utility protocol already suffices for efficient communication. In particular, in a game with possible observations and a…

cs.AI2025

Observation Interference in Partially Observable Assistance Games

Scott Emmons, Caspar Oesterheld, Vincent Conitzer +1

We study partially observable assistance games (POAGs), a model of the human-AI value alignment problem which allows the human and the AI assistant to have partial observations. Mo…

cs.LG2025

Obfuscated Activations Bypass LLM Latent-Space Defenses

Luke Bailey, Alex Serrano, Abhay Sheshadri +7

Recent latent-space monitoring techniques have shown promise as defenses against LLM attacks. These defenses act as scanners that seek to detect harmful activations before they lea…

cs.LG2025

ALMANACS: A Simulatability Benchmark for Language Model Explainability

Edmund Mills, Shiye Su, Stuart Russell +1

How do we measure the efficacy of language model explainability methods? While many explainability methods have been developed, they are typically evaluated on bespoke tasks, preve…

cs.GT2024

The Partially Observable Off-Switch Game

Andrew Garber, Rohan Subramani, Linus Luu +3

A wide variety of goals could cause an AI to disable its off switch because "you can't fetch the coffee if you're dead" (Russell 2019). Prior theoretical work on this shutdown prob…

cs.LG2024

When Your AIs Deceive You: Challenges of Partial Observability in Reinforcement Learning from Human Feedback

Leon Lang, Davis Foote, Stuart Russell +3

Past analyses of reinforcement learning from human feedback (RLHF) assume that the human evaluators fully observe the environment. What happens when human feedback is based only on…