Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
OpenACMv2: An Accuracy-Constrained Co-Optimization Framework for Approximate DCiM
Yiqi Zhou, Yue Yuan, Yikai Wang +8
Digital Compute-in-Memory (DCiM) accelerates neural networks by reducing data movement. Approximate DCiM can further improve power-performance-area (PPA), but demands accuracy-cons…
cs.LG2025
X-MoE: Enabling Scalable Training for Emerging Mixture-of-Experts Architectures on HPC Platforms
Yueming Yuan, Ahan Gupta, Jianping Li +3
Emerging expert-specialized Mixture-of-Experts (MoE) architectures, such as DeepSeek-MoE, deliver strong model quality through fine-grained expert segmentation and large top-k rout…
cs.LG2025
MiLo: Efficient Quantized MoE Inference with Mixture of Low-Rank Compensators
Beichen Huang, Yueming Yuan, Zelei Shao +1
A critical approach for efficiently deploying Mixture-of-Experts (MoE) models with massive parameters is quantization. However, state-of-the-art MoE models suffer from non-negligib…