collaborators

6 papers

cs.LG2026

Modular Pretraining Enables Access Control

Ethan Roland, Murat Cubuktepe, Erick Martinez +8

AI developers face a dual-use dilemma. An AI capability that helps one user cure a disease can help another synthesize one. This dilemma could be resolved with access control, limi…

cs.LG2026

Endogenous Resistance to Activation Steering in Language Models

Alex McKenzie, Keenan Pepper, Stijn Servaes +6

Large language models can recover mid-generation from task-misaligned activation steering, producing explicit verbal restarts (e.g., ``wait, that's not right'') and continuing on-t…

cs.CL2026

Learning Self-Interpretation from Interpretability Artifacts: Training Lightweight Adapters on Vector-Label Pairs

Keenan Pepper, Alex McKenzie, Florin Pop +6

Self-interpretation methods prompt language models to describe their own internal states, but remain unreliable due to hyperparameter sensitivity. We show that training lightweight…

cs.AI2024

Towards Safe and Honest AI Agents with Neural Self-Other Overlap

Marc Carauleanu, Michael Vaiana, Judd Rosenblatt +2

As AI systems increasingly make critical decisions, deceptive AI poses a significant challenge to trust and safety. We present Self-Other Overlap (SOO) fine-tuning, a promising app…

cs.LG2024

Unexpected Benefits of Self-Modeling in Neural Systems

Vickram N. Premakumar, Michael Vaiana, Florin Pop +4

Self-models have been a topic of great interest for decades in studies of human cognition and more recently in machine learning. Yet what benefits do self-models confer? Here we sh…

cs.CL2024

Rethinking harmless refusals when fine-tuning foundation models

Florin Pop, Judd Rosenblatt, Diogo Schwerz de Lucena +1

In this paper, we investigate the degree to which fine-tuning in Large Language Models (LLMs) effectively mitigates versus merely conceals undesirable behavior. Through the lens of…