5 papers · 1 filter
SAE Interventions are Unreliable: Post-Intervention Recovery of Suppressed Behavior
Mingyue Cui, Linghui Shen, Xingyi Yang
Sparse Autoencoders (SAEs) decompose residual-stream activations into interpretable features. Recent latent-space defenses increasingly rely on these decompositions, assuming that…
DirMoE: Dirichlet-routed Mixture of Experts
Amirhossein Vahidi, Hesam Asadollahzadeh, Navid Akhavan Attar +4
Mixture-of-Experts (MoE) models have demonstrated exceptional performance in large-scale language models. Existing routers typically rely on non-differentiable Top-+Softmax, lim…
Don't Forget the Nonlinearity: Unlocking Activation Functions in Efficient Fine-Tuning
Bo Yin, Xingyi Yang, Xinchao Wang
Existing parameter-efficient fine-tuning (PEFT) methods primarily adapt weight matrices while keeping activation functions fixed. We introduce \textbf{NoRA}, the first PEFT framewo…
Mixture of Experts Made Intrinsically Interpretable
Xingyi Yang, Constantin Venhoff, Ashkan Khakzar +4
Neurons in large language models often exhibit \emph{polysemanticity}, simultaneously encoding multiple unrelated concepts and obscuring interpretability. Instead of relying on pos…
AdvAnchor: Enhancing Diffusion Model Unlearning with Adversarial Anchors
Mengnan Zhao, Lihe Zhang, Xingyi Yang +2
Security concerns surrounding text-to-image diffusion models have driven researchers to unlearn inappropriate concepts through fine-tuning. Recent fine-tuning methods typically ali…