11 papers
Distribution-Specific Curvature Control with Finite-Sample Guarantees for Open-Weight Safety
Domenic Rosati, Ali Dadsetan, Hong Huang +5
A short fine-tuning run can undo the safety guards of an open-weight model---retraining a refusal-trained assistant to aid weapons development or produce hate speech. Preventing su…
Multi-Granular Node Pruning for Causal Circuit Discovery
Muhammad Umair Haider, Hammad Rizwan, Hassan Sajjad +1
Circuit discovery aims to identify minimal subnetworks that are responsible for specific behaviors in large language models (LLMs). Existing approaches primarily rely on iterative…
Vector Quantized Latent Concepts: A Scalable Alternative to Clustering-Based Concept Discovery
Xuemin Yu, Ankur Garg, Samira Ebrahimi Kahou +1
Large language models (LLMs) encode rich semantic information in their hidden states, yet it remains difficult to understand what information these internal representations capture…
Neurons Speak in Ranges: Breaking Free from Discrete Neuronal Attribution
Muhammad Umair Haider, Hammad Rizwan, Hassan Sajjad +2
Pervasive polysemanticity in large language models (LLMs) undermines discrete neuron-concept attribution, posing a significant challenge for model interpretation and control. We sy…
Limits of Convergence-Rate Control for Open-Weight Safety
Domenic Rosati, Xijie Zeng, Hong Huang +4
Open-weight foundation models can be fine-tuned for harmful purposes after release, yet no existing training resistance methods provide theoretical guarantees. Treating these inter…
Dependency Parsing is More Parameter-Efficient with Normalization
Paolo Gajo, Domenic Rosati, Hassan Sajjad +1
Dependency parsing is the task of inferring natural language structure, often approached by modeling word interactions via attention through biaffine scoring. This mechanism works…