257 citations · 494 across the 20 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2024
Exploring Activation Patterns of Parameters in Language Models
Yudong Wang, Damai Dai, Zhifang Sui
Most work treats large language models as black boxes without in-depth understanding of their internal working mechanism. In order to explain the internal representations of LLMs,…
cs.LG2022
StableMoE: Stable Routing Strategy for Mixture of Experts
Damai Dai, Li Dong, Shuming Ma +4
The Mixture-of-Experts (MoE) technique can scale up the model size of Transformers with an affordable computational overhead. We point out that existing learning-to-route MoE metho…