3 citations · 3 across the 2 of their papers we have counts for
3 papers
Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
Adly Templeton, Tom Conerly, Jonathan Marcus +23
We demonstrate that sparse autoencoders can extract interpretable features from Claude 3 Sonnet, a production-scale language model, addressing the open question of whether dictiona…
Must Read: A Comprehensive Survey of Computational Persuasion
Nimet Beyza Bozdag, Shuhaib Mehri, Xiaocheng Yang +7
Persuasion is a fundamental aspect of communication, influencing decision-making across diverse contexts, from everyday conversations to high-stakes scenarios such as politics, mar…
Stress-Testing Model Specs Reveals Character Differences among Language Models
Jifan Zhang, Henry Sleight, Andi Peng +2
Large language models (LLMs) are increasingly trained from AI constitutions and model specifications that establish behavioral guidelines and ethical principles. However, these spe…