activity
20242026
collaborators
Showing cs.LGShow all

9 papers · 1 filter

cs.LG2026

Decoy Direction Optimization: A Post-Hoc Defense Against LLM Abliteration

Aashiq Muhamed, Mona T. Diab, Virginia Smith

Safety guardrails in open-weight language models can be readily bypassed using Refusal Feature Ablation (RFA), a technique that identifies and projects out a linear refusal directi…

cs.LG2026

Pick Your Poison: Learning to Select Poison Sets for Stronger LLM Backdoor Attacks

Aashiq Muhamed, Mona T. Diab, Virginia Smith +2

Backdoor poisoning attacks add poisoned examples to otherwise-clean finetuning data, pairing a trigger with a target behavior that the model learns to produce when the trigger appe…

cs.LG2026

RAPTOR: Role-Aware Private Training for Mixture-of-Experts

Duc Dm, Khai Le-Duc, Nguyen Do +18

Differentially private (DP) fine-tuning methods treat sparse Mixture-of-Experts (MoE) models as a single dense block, ignoring that shared layers see all data while experts only se…

cs.LG2026

Pando: Do Interpretability Methods Work When Models Won't Explain Themselves?

Ziqian Zhong, Aashiq Muhamed, Mona T. Diab +2

Mechanistic interpretability is often motivated for alignment auditing, where a model's verbal explanations can be absent, incomplete, or misleading. Yet many evaluations do not co…

cs.LG2026

DSPA: Dynamic SAE Steering for Data-Efficient Preference Alignment

James Wedgwood, Aashiq Muhamed, Mona T. Diab +1

Preference alignment is usually achieved by weight-updating training on preference data, which adds substantial alignment-stage compute and provides limited mechanistic visibility.…

cs.LG2025

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs

Xiangchen Song, Aashiq Muhamed, Yujia Zheng +5

Sparse Autoencoders (SAEs) are a prominent tool in mechanistic interpretability (MI) for decomposing neural network activations into interpretable features. However, the aspiration…