2 citations · 3 across the 5 of their papers we have counts for
6 papers · 1 filter
FlashSchNet: Fast and Accurate Coarse-Grained Neural Network Molecular Dynamics
Pingzhi Li, Hongxuan Li, Zirui Liu +2
Graph neural network (GNN) potentials such as SchNet improve the accuracy and transferability of molecular dynamics (MD) simulation by learning many-body interactions, but remain s…
Towards Building Non-Fine-Tunable Foundation Models
Ziyao Wang, Nizhang Li, Pingzhi Li +3
Open-sourcing foundation models (FMs) enables broad reuse but also exposes model trainers to economic and safety risks from unrestricted downstream fine-tuning. We address this pro…
Occult: Optimizing Collaborative Communication across Experts for Accelerated Parallel MoE Training and Inference
Shuqing Luo, Pingzhi Li, Jie Peng +7
Mixture-of-experts (MoE) architectures could achieve impressive computational efficiency with expert parallelism, which relies heavily on all-to-all communication across devices. U…
Glider: Global and Local Instruction-Driven Expert Router
Pingzhi Li, Prateek Yadav, Jaehong Yoon +4
The availability of performant pre-trained models has led to a proliferation of fine-tuned expert models that are specialized to particular domains. This has enabled the creation o…
QuantMoE-Bench: Examining Post-Training Quantization for Mixture-of-Experts
Pingzhi Li, Xiaolong Jin, Zhen Tan +2
Mixture-of-Experts (MoE) is a promising way to scale up the learning capacity of large language models. It increases the number of parameters while keeping FLOPs nearly constant du…
Model-GLUE: Democratized LLM Scaling for A Large Model Zoo in the Wild
Xinyu Zhao, Guoheng Sun, Ruisi Cai +13
As Large Language Models (LLMs) excel across tasks and specialized domains, scaling LLMs based on existing models has garnered significant attention, which faces the challenge of d…