6 papers
Computationally Efficient Collaborative Communication Via Regularity-Based Coarsening
Mark Bedaywi, Scott Emmons, Nika Haghtalab +1
Our results show that the existence of a short high-utility protocol already suffices for efficient communication. In particular, in a game with possible observations and a…
Observation Interference in Partially Observable Assistance Games
Scott Emmons, Caspar Oesterheld, Vincent Conitzer +1
We study partially observable assistance games (POAGs), a model of the human-AI value alignment problem which allows the human and the AI assistant to have partial observations. Mo…
Obfuscated Activations Bypass LLM Latent-Space Defenses
Luke Bailey, Alex Serrano, Abhay Sheshadri +7
Recent latent-space monitoring techniques have shown promise as defenses against LLM attacks. These defenses act as scanners that seek to detect harmful activations before they lea…
ALMANACS: A Simulatability Benchmark for Language Model Explainability
Edmund Mills, Shiye Su, Stuart Russell +1
How do we measure the efficacy of language model explainability methods? While many explainability methods have been developed, they are typically evaluated on bespoke tasks, preve…
The Partially Observable Off-Switch Game
Andrew Garber, Rohan Subramani, Linus Luu +3
A wide variety of goals could cause an AI to disable its off switch because "you can't fetch the coffee if you're dead" (Russell 2019). Prior theoretical work on this shutdown prob…
When Your AIs Deceive You: Challenges of Partial Observability in Reinforcement Learning from Human Feedback
Leon Lang, Davis Foote, Stuart Russell +3
Past analyses of reinforcement learning from human feedback (RLHF) assume that the human evaluators fully observe the environment. What happens when human feedback is based only on…