2 papers
cs.LG2024
MoNTA: Accelerating Mixture-of-Experts Training with Network-Traffc-Aware Parallel Optimization
Jingming Guo, Yan Liu, Yu Meng +4
The Mixture of Experts (MoE) is an advanced model architecture in the industry that combines multiple specialized expert models from various domains into a single supermodel. This…
cs.CV2022
DenseShift: Towards Accurate and Efficient Low-Bit Power-of-Two Quantization
Xinlin Li, Bang Liu, Rui Heng Yang +3
Efficiently deploying deep neural networks on low-resource edge devices is challenging due to their ever-increasing resource requirements. To address this issue, researchers have p…