Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
BEAM: Binary Expert Activation Masking for Dynamic Routing in MoE
Juntong Wu, Jialiang Cheng, Qishen Yin +5
Mixture-of-Experts (MoE) architectures enhance the efficiency of large language models by activating only a subset of experts per token. However, standard MoE employs a fixed Top-K…
cs.AI2024
MultiDelete for Multimodal Machine Unlearning
Jiali Cheng, Hadi Amiri
Machine Unlearning removes specific knowledge about training data samples from an already trained model. It has significant practical benefits, such as purging private, inaccurate,…