5 papers
Subliminal Learning: Language models transmit behavioral traits via hidden signals in data
Alex Cloud, Minh Le, James Chua +5
We study subliminal learning, a surprising phenomenon where language models transmit behavioral traits via semantically unrelated data. In our main experiments, a "teacher" model w…
Obfuscated Activations Bypass LLM Latent-Space Defenses
Luke Bailey, Alex Serrano, Abhay Sheshadri +7
Recent latent-space monitoring techniques have shown promise as defenses against LLM attacks. These defenses act as scanners that seek to detect harmful activations before they lea…
Towards a Law of Iterated Expectations for Heuristic Estimators
Paul Christiano, Jacob Hilton, Andrea Lincoln +2
Christiano et al. (2022) define a *heuristic estimator* to be a hypothetical algorithm that estimates the values of mathematical expressions from arguments. In brief, a heuristic e…
Estimating the Probabilities of Rare Outputs in Language Models
Gabriel Wu, Jacob Hilton
We consider the problem of low probability estimation: given a machine learning model and a formally-specified input distribution, how can we estimate the probability of a binary p…
Backdoor defense, learnability and obfuscation
Paul Christiano, Jacob Hilton, Victor Lecomte +1
We introduce a formal notion of defendability against backdoors using a game between an attacker and a defender. In this game, the attacker modifies a function to behave differentl…