11 papers
Mind Viruses: Self-Propagating Ideas in Multi-Agent LLM Systems
Vassilis Papadopoulos, McNair Shah, Sam Zimmerman +1
AI agents are becoming more autonomous and increasingly interconnected, exposing them to new emergent risks arising from agent-to-agent interaction. One such risk is the spread of…
Verbalizable Representations Form a Global Workspace in Language Models
Wes Gurnee, Nicholas Sofroniew, Adam Pearce +13
Out of everything the human brain processes, only a small fraction is consciously accessible, in the sense of being available for verbal report, deliberate control, and flexible re…
The Value Axis: Language Models Encode Whether They're on the Right Track
Nick Jiang, Isaac Kauvar, Jack Lindsey
We investigate whether language models internally track the value of their current trajectory, defined as the likelihood that their ongoing strategy will achieve their goals. Using…
Mechanisms of Introspective Awareness
Uzay Macar, Li Yang, Atticus Wang +3
Recent work has shown that LLMs can sometimes detect when steering vectors are injected into their residual stream and identify the injected concept -- a phenomenon termed "introsp…
SAE-StatSteer: Statistical Consensus Feature Selection for Optimization-Free Activation Steering of Large Language Models
Oshayer Siddique, J. M Areeb Uzair Alam, Md Jobayer Rahman Rafy +3
Activation steering adds a residual-stream direction at inference time, providing lightweight behavioral control without fine-tuning. Sparse autoencoders (SAEs) can make such inter…
Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
Adly Templeton, Tom Conerly, Jonathan Marcus +23
We demonstrate that sparse autoencoders can extract interpretable features from Claude 3 Sonnet, a production-scale language model, addressing the open question of whether dictiona…