6 papers
Verbalizable Representations Form a Global Workspace in Language Models
Wes Gurnee, Nicholas Sofroniew, Adam Pearce +13
Out of everything the human brain processes, only a small fraction is consciously accessible, in the sense of being available for verbal report, deliberate control, and flexible re…
Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers
Adam Karvonen, James Chua, Clément Dumas +8
Large language model (LLM) activations are notoriously difficult to understand, with most existing techniques using complex, specialized methods for interpreting them. Recent work…
Scaling Laws For Scalable Oversight
Joshua Engels, David D. Baek, Subhash Kantamneni +1
Scalable oversight, the process by which weaker AI systems supervise stronger ones, has been proposed as a key strategy to control future superintelligent systems. However, it is s…
The Singapore Consensus on Global AI Safety Research Priorities
Yoshua Bengio, Tegan Maharaj, Luke Ong +84
Rapidly improving AI capabilities and autonomy hold significant promise of transformation, but are also driving vigorous debate on how to ensure that AI is safe, i.e., trustworthy,…
Are Sparse Autoencoders Useful? A Case Study in Sparse Probing
Subhash Kantamneni, Joshua Engels, Senthooran Rajamanoharan +2
Sparse autoencoders (SAEs) are a popular method for interpreting concepts represented in large language model (LLM) activations. However, there is a lack of evidence regarding the…
Language Models Use Trigonometry to Do Addition
Subhash Kantamneni, Max Tegmark
Mathematical reasoning is an increasingly important indicator of large language model (LLM) capabilities, yet we lack understanding of how LLMs process even simple mathematical tas…