From the 3 of 19 linked papers with an AI index.
1 citations · 1 across the 12 of their papers we have counts for
3 papers · 1 filter
BEAM: Binary Expert Activation Masking for Dynamic Routing in MoE
Juntong Wu, Jialiang Cheng, Qishen Yin +5
Mixture-of-Experts (MoE) architectures enhance the efficiency of large language models by activating only a subset of experts per token. However, standard MoE employs a fixed Top-K…
SlimGPT: Layer-wise Structured Pruning for Large Language Models
Gui Ling, Ziyang Wang, Yuliang Yan +1
Large language models (LLMs) have garnered significant attention for their remarkable capabilities across various domains, whose vast parameter scales present challenges for practi…
ChemLLM: A Chemical Large Language Model
Di Zhang, Wei Liu, Qian Tan +12
Large language models (LLMs) have made impressive progress in chemistry applications. However, the community lacks an LLM specifically designed for chemistry. The main challenges a…