4 papers
Exploitation Without Deception: Dark Triad Feature Steering Reveals Separable Antisocial Circuits in Language Models
Cameron Berg, Roshni Lulla
We use sparse autoencoder (SAE) feature steering to amplify Dark Triad personality traits (Machiavellianism, narcissism, and psychopathy) in Llama-3.3-70B-Instruct and evaluate the…
Hidden Coalitions in Multi-Agent AI: A Spectral Diagnostic from Internal Representations
Cameron Berg, Susan L. Schneider, Mark M. Bailey
Collections of interacting AI agents can form coalitions, creating emergent group-level organization that is critical for AI safety and alignment. However, observing agent behavior…
Large Language Models Report Subjective Experience Under Self-Referential Processing
Cameron Berg, Diogo de Lucena, Judd Rosenblatt
Large language models sometimes produce structured, first-person descriptions that explicitly reference awareness or subjective experience. To better understand this behavior, we i…
Towards Safe and Honest AI Agents with Neural Self-Other Overlap
Marc Carauleanu, Michael Vaiana, Judd Rosenblatt +2
As AI systems increasingly make critical decisions, deceptive AI poses a significant challenge to trust and safety. We present Self-Other Overlap (SOO) fine-tuning, a promising app…