6 papers
Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
Adly Templeton, Tom Conerly, Jonathan Marcus +23
We demonstrate that sparse autoencoders can extract interpretable features from Claude 3 Sonnet, a production-scale language model, addressing the open question of whether dictiona…
Are Large Language Models Sensitive to the Motives Behind Communication?
Addison J. Wu, Ryan Liu, Kerem Oktar +2
Human communication is motivated: people speak, write, and create content with a particular communicative intent in mind. As a result, information that large language models (LLMs)…
Rational Metareasoning for Large Language Models
C. Nicolò De Sabbata, Theodore R. Sumers, Badr AlKhamissi +2
Being prompted to engage in reasoning has emerged as a core technique for using large language models (LLMs), deploying additional inference-time compute to improve task performanc…
Representational Alignment Supports Effective Machine Teaching
Ilia Sucholutsky, Katherine M. Collins, Maya Malaviya +11
A good teacher should not only be knowledgeable, but should also be able to communicate in a way that the student understands -- to share the student's representation of the world.…
Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming
Mrinank Sharma, Meg Tong, Jesse Mu +40
Large language models (LLMs) are vulnerable to universal jailbreaks-prompting strategies that systematically bypass model safeguards and enable users to carry out harmful processes…
Clio: Privacy-Preserving Insights into Real-World AI Use
Alex Tamkin, Miles McCain, Kunal Handa +18
How are AI assistants being used in the real world? While model providers in theory have a window into this impact via their users' data, both privacy concerns and practical challe…