activity
20242026
collaborators
Showing cs.LGShow all

5 papers · 1 filter

cs.LG2026

SAE Interventions are Unreliable: Post-Intervention Recovery of Suppressed Behavior

Mingyue Cui, Linghui Shen, Xingyi Yang

Sparse Autoencoders (SAEs) decompose residual-stream activations into interpretable features. Recent latent-space defenses increasingly rely on these decompositions, assuming that…

cs.LG2026

DirMoE: Dirichlet-routed Mixture of Experts

Amirhossein Vahidi, Hesam Asadollahzadeh, Navid Akhavan Attar +4

Mixture-of-Experts (MoE) models have demonstrated exceptional performance in large-scale language models. Existing routers typically rely on non-differentiable Top-+Softmax, lim…

cs.LG2025

Don't Forget the Nonlinearity: Unlocking Activation Functions in Efficient Fine-Tuning

Bo Yin, Xingyi Yang, Xinchao Wang

Existing parameter-efficient fine-tuning (PEFT) methods primarily adapt weight matrices while keeping activation functions fixed. We introduce \textbf{NoRA}, the first PEFT framewo…

cs.LG2025

Mixture of Experts Made Intrinsically Interpretable

Xingyi Yang, Constantin Venhoff, Ashkan Khakzar +4

Neurons in large language models often exhibit \emph{polysemanticity}, simultaneously encoding multiple unrelated concepts and obscuring interpretability. Instead of relying on pos…

cs.LG2024

AdvAnchor: Enhancing Diffusion Model Unlearning with Adversarial Anchors

Mengnan Zhao, Lihe Zhang, Xingyi Yang +2

Security concerns surrounding text-to-image diffusion models have driven researchers to unlearn inappropriate concepts through fine-tuning. Recent fine-tuning methods typically ali…