2 papers
cs.DC2025
MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism
Zheng Zhang, Donglin Yang, Yaqi Xia +4
Recently, Mixture-of-Experts (MoE) has become one of the most popular techniques to scale pre-trained models to extraordinarily large sizes. Dynamic activation of experts allows fo…
cs.DC2025
MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators
Zheng Zhang, Donglin Yang, Xiaobo Zhou +1
Operator fusion, a key technique to improve data locality and alleviate GPU memory bandwidth pressure, often fails to extend to the fusion of multiple compute-intensive operators d…