9 papers · 1 filter
Decoy Direction Optimization: A Post-Hoc Defense Against LLM Abliteration
Aashiq Muhamed, Mona T. Diab, Virginia Smith
Safety guardrails in open-weight language models can be readily bypassed using Refusal Feature Ablation (RFA), a technique that identifies and projects out a linear refusal directi…
Pick Your Poison: Learning to Select Poison Sets for Stronger LLM Backdoor Attacks
Aashiq Muhamed, Mona T. Diab, Virginia Smith +2
Backdoor poisoning attacks add poisoned examples to otherwise-clean finetuning data, pairing a trigger with a target behavior that the model learns to produce when the trigger appe…
RAPTOR: Role-Aware Private Training for Mixture-of-Experts
Duc Dm, Khai Le-Duc, Nguyen Do +18
Differentially private (DP) fine-tuning methods treat sparse Mixture-of-Experts (MoE) models as a single dense block, ignoring that shared layers see all data while experts only se…
Pando: Do Interpretability Methods Work When Models Won't Explain Themselves?
Ziqian Zhong, Aashiq Muhamed, Mona T. Diab +2
Mechanistic interpretability is often motivated for alignment auditing, where a model's verbal explanations can be absent, incomplete, or misleading. Yet many evaluations do not co…
DSPA: Dynamic SAE Steering for Data-Efficient Preference Alignment
James Wedgwood, Aashiq Muhamed, Mona T. Diab +1
Preference alignment is usually achieved by weight-updating training on preference data, which adds substantial alignment-stage compute and provides limited mechanistic visibility.…
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs
Xiangchen Song, Aashiq Muhamed, Yujia Zheng +5
Sparse Autoencoders (SAEs) are a prominent tool in mechanistic interpretability (MI) for decomposing neural network activations into interpretable features. However, the aspiration…