activity
20242026
collaborators

6 papers

cs.AI2026

Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet

Adly Templeton, Tom Conerly, Jonathan Marcus +23

We demonstrate that sparse autoencoders can extract interpretable features from Claude 3 Sonnet, a production-scale language model, addressing the open question of whether dictiona…

cs.CL2026

Are Large Language Models Sensitive to the Motives Behind Communication?

Addison J. Wu, Ryan Liu, Kerem Oktar +2

Human communication is motivated: people speak, write, and create content with a particular communicative intent in mind. As a result, information that large language models (LLMs)…

cs.CL2025

Rational Metareasoning for Large Language Models

C. Nicolò De Sabbata, Theodore R. Sumers, Badr AlKhamissi +2

Being prompted to engage in reasoning has emerged as a core technique for using large language models (LLMs), deploying additional inference-time compute to improve task performanc…

cs.LG2025

Representational Alignment Supports Effective Machine Teaching

Ilia Sucholutsky, Katherine M. Collins, Maya Malaviya +11

A good teacher should not only be knowledgeable, but should also be able to communicate in a way that the student understands -- to share the student's representation of the world.…

cs.CL2025

Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming

Mrinank Sharma, Meg Tong, Jesse Mu +40

Large language models (LLMs) are vulnerable to universal jailbreaks-prompting strategies that systematically bypass model safeguards and enable users to carry out harmful processes…

cs.CY2024

Clio: Privacy-Preserving Insights into Real-World AI Use

Alex Tamkin, Miles McCain, Kunal Handa +18

How are AI assistants being used in the real world? While model providers in theory have a window into this impact via their users' data, both privacy concerns and practical challe…