5 papers
Subliminal Learning: Language models transmit behavioral traits via hidden signals in data
Alex Cloud, Minh Le, James Chua +5
We study subliminal learning, a surprising phenomenon where language models transmit behavioral traits via semantically unrelated data. In our main experiments, a "teacher" model w…
Obfuscated Activations Bypass LLM Latent-Space Defenses
Luke Bailey, Alex Serrano, Abhay Sheshadri +7
Recent latent-space monitoring techniques have shown promise as defenses against LLM attacks. These defenses act as scanners that seek to detect harmful activations before they lea…
Estimating the Probabilities of Rare Outputs in Language Models
Gabriel Wu, Jacob Hilton
We consider the problem of low probability estimation: given a machine learning model and a formally-specified input distribution, how can we estimate the probability of a binary p…
Backdoor defense, learnability and obfuscation
Paul Christiano, Jacob Hilton, Victor Lecomte +1
We introduce a formal notion of defendability against backdoors using a game between an attacker and a defender. In this game, the attacker modifies a function to behave differentl…
Towards a Law of Iterated Expectations for Heuristic Estimators
Paul Christiano, Jacob Hilton, Andrea Lincoln +2
Christiano et al. (2022) define a *heuristic estimator* to be a hypothetical algorithm that estimates the values of mathematical expressions from arguments. In brief, a heuristic e…