10 papers
Building Production-Ready Probes For Gemini
János Kramár, Joshua Engels, Zheng Wang +4
Frontier language model capabilities are improving rapidly. We thus need stronger mitigations against bad actors misusing increasingly powerful systems. Prior work has shown that a…
Simple Mechanistic Explanations for Out-Of-Context Reasoning
Atticus Wang, Joshua Engels, Oliver Clive-Griffin +2
Out-of-context reasoning (OOCR) is a phenomenon in which fine-tuned LLMs exhibit surprisingly deep out-of-distribution generalization. Rather than learning shallow heuristics, they…
The Singapore Consensus on Global AI Safety Research Priorities
Yoshua Bengio, Tegan Maharaj, Luke Ong +84
Rapidly improving AI capabilities and autonomy hold significant promise of transformation, but are also driving vigorous debate on how to ensure that AI is safe, i.e., trustworthy,…
Dense SAE Latents Are Features, Not Bugs
Xiaoqing Sun, Alessandro Stolfo, Joshua Engels +4
Sparse autoencoders (SAEs) are designed to extract interpretable features from language models by enforcing a sparsity constraint. Ideally, training an SAE would yield latents that…
Scaling Laws For Scalable Oversight
Joshua Engels, David D. Baek, Subhash Kantamneni +1
Scalable oversight, the process by which weaker AI systems supervise stronger ones, has been proposed as a key strategy to control future superintelligent systems. However, it is s…
Are Sparse Autoencoders Useful? A Case Study in Sparse Probing
Subhash Kantamneni, Joshua Engels, Senthooran Rajamanoharan +2
Sparse autoencoders (SAEs) are a popular method for interpreting concepts represented in large language model (LLM) activations. However, there is a lack of evidence regarding the…